anydoc routes eight office formats through one Markdown serializer, and PDF around it
Watch on TikTok
The video claims that every format anydoc supports parses into one shared document model and renders through one serializer, so a table-escaping fix for docx also fixes rtf and odt. That claim holds for the eight office formats and does not hold for PDF, which the source code routes around the document model entirely. The frames show github.com/firecrawl/anydoc: a README header with crates.io, npm and PyPI badges all reading v0.1.3, an MIT license badge, the repo page at 1.5k stars and 63 forks, the Supported formats table, the benchmark table against six other converters, and the "How it works" ASCII diagram. The last two frames show that diagram, where PDF branches off to a separate library, while the voiceover says every format goes through the model. The video also skips what anydoc does when a PDF is scanned, which at the version on screen was to drop those pages without saying so.
What the video shows, and what has moved since
The video went up on 2026-08-05, two days after the repository was created on 2026-08-03. The frames pin it to v0.1.3 across all three registries, with 1.5k stars, 63 forks, 6 open issues and 2 open pull requests.
Checked on 2026-09-29, the same repository reads 22,267 stars, 1,404 forks and 101 open issues, with v0.2.4 released on 2026-08-27 and the last push on 2026-08-28. Star counts move, so read that number as one day's snapshot rather than a property of the project. The published packages are the crate anydoc, the npm package @firecrawl/anydoc, and the PyPI package firecrawl-anydoc, all at 0.2.4. A fourth package, @firecrawl/anydoc-wasm, plus a browser demo that converts files locally, arrived after the video. The project is spelled anydoc in lowercase everywhere in the README, it is maintained by Firecrawl, and the GitHub account tomsideguide has written the large majority of commits.
The architecture claim is true for office formats and false for PDF
The README bullet the video paraphrases says each format "parses into a shared document model and renders through a single Markdown serializer." The source supports that for doc, docx, ppt, pptx, xls, xlsx, xlsb, odt, ods, odp, rtf, epub and csv. Each has a parser under src/formats/, the model lives in src/model/, and one serializer lives in src/render/markdown/ with escape.rs and table.rs as separate files inside it. A change to escape.rs does reach all of those formats in one commit.
PDF takes a different path. In src/lib.rs, to_markdown_bytes short-circuits before the model is ever built:
// PDFs convert to Markdown directly (pdf-inspector) without passing
// through the document model.
if format == Format::Pdf {
return formats::pdf::to_markdown(bytes);
}
The public to_document function carries a doc comment saying it is "Unsupported for Format::Pdf: PDF conversion produces Markdown directly and has no document-model form." The README's own diagram draws PDF → pdf-inspector → Markdown directly as a branch off the main flow. So a table-escaping fix in the shared serializer reaches rtf and odt as the video says, and it does not reach PDF output. If you are picking anydoc because one fix covers everything, scope that expectation to the office formats.
Format detection reads bytes, except for CSV
This part of the video checks out, with one exception it does not mention. In src/formats/detect.rs, from_bytes tests for the RTF open group {\rtf, the OLE compound file signature D0 CF 11 E0 A1 B1 1A E1, the ZIP local file header PK\x03\x04, and the %PDF- header anywhere in the first 1024 bytes. OLE files are then classified by stream name, looking for WordDocument, PowerPoint Document, Workbook or Book. ZIP packages are classified by the mimetype part for ODF and EPUB, or by the OPC content type of the part the package-level relationship designates for the OOXML formats. None of that reads the filename.
CSV is the exception, and the module documentation states it: "Plain-text formats (CSV) carry no signature and are never detected; callers fall back to the file extension." to_markdown calls Format::from_bytes first and falls back to Format::from_path. Hand CSV bytes to to_markdown_bytes with no format argument and you get Unsupported back. The video's line that mislabeled files still convert is true for seven of the eight format families and false for the eighth.
Where PDF-to-Markdown actually breaks
A PDF stores placed glyphs with coordinates. It does not store paragraphs, columns or table cells unless someone tagged it, and most PDFs in the wild are untagged. Three cases cause most of the damage. Scanned pages carry images with no text layer at all. Multi-column layouts require reading order to be reconstructed from geometry, and getting it wrong interleaves two columns line by line. Tables drawn without ruling lines give the extractor nothing but whitespace gaps to infer column boundaries from.
anydoc is direct about the first case and quiet about the other two. The README says it performs no OCR and lists NeedsOcr among its error variants. The v0.2.4 release notes, dated 2026-08-27, say that before that release "a PDF with scanned or image-only pages used to convert with those pages silently missing, and a fully scanned one failed as unsupported." The version on screen in the video, v0.1.3, had exactly that silent-drop behavior, which is the single most useful thing the video could have told a viewer and did not. The in-library fix is to fail loudly with the page numbers. The fix for the document is --ocr hosted, which uploads the whole file to Firecrawl Parse, the hosted service run by the same company. No signup is required for low volume, and an API key raises the limits.
For multi-column layout and unruled tables, anydoc delegates to pdf-inspector, its sibling Rust library, and takes the speed side of the tradeoff. The README states the position plainly: "Pure Rust, no ML models, no external services." Docling sits on the other side of that tradeoff, shipping layout and table-structure models that reconstruct reading order and cell boundaries, at a median of 513.6ms per document in anydoc's own benchmark against anydoc's 4.4ms. Both are defensible. They are not substitutes for each other on a stack of scanned invoices.
What the benchmark proves, and what it does not
The README benchmark runs 100 real-world documents across fourteen formats against six other converters. anydoc covers 14 of 14 formats at a median 4.4ms and an overall score of 81. markitdown scores 65 over 6 formats, unstructured 63 over 8, docling 57 over 4, pandoc 56 over 5, libreoffice 40 over 12, and mammoth 70 over docx alone. The per-format table is the fair comparison, and anydoc leads every row in it.
Some of those gaps are scope rather than quality. Pandoc is a universal markup converter with dozens of input formats, and binary office formats have never been its centre of gravity. mammoth only ever claimed docx. markitdown is Microsoft's and aims at breadth of file type, including audio and images, rather than fidelity on a .ppt from 2003.
Four caveats sit in the README's own fine print, and they matter if you are going to cite these numbers. The judge is an LLM, Claude Sonnet 5, scoring blind pairs with outputs swapped to cancel position bias across 482 verdicts. The ground truth is the document's first six pages rendered to images by LibreOffice, so nothing past page six is judged and the yardstick is produced by one of the seven tools being scored. The corpus is not redistributable and is not in the repository, so nobody outside Firecrawl can reproduce the run. Timings exclude process spawn for anydoc and the Python libraries and include it for the CLI tools. The video's frame shows a 4.7ms median where the README now shows 4.4ms, which at least indicates the harness gets re-run rather than frozen at launch.
What the shared model costs
The document model is the ceiling on fidelity for every format that passes through it. src/model/ carries blocks, inlines, tables, lists, links, styles and assets. Anything a parser can see that the model has no shape for is gone before the serializer runs, and no serializer work recovers it. PDF is the proof that the ceiling is real. It did not fit the model, so it got its own path and its own separate quality characteristics.
The second cost is blast radius. The property the video sells, that a docx table fix lands in rtf and odt, runs in both directions. A regression in src/render/markdown/escape.rs breaks thirteen formats in one commit. The repository answers that with a committed fixture corpus under tests/fixtures/ that is snapshot-tested, tests/robustness.rs mutation-testing every fixture, and cargo-fuzz targets per format under fuzz/. Those snapshots are what stands between one serializer change and a quiet quality drop across the whole matrix, so their coverage is the thing to inspect if you are evaluating this for a production pipeline.
Key Takeaways
- The correct repository is github.com/firecrawl/anydoc, MIT licensed, created 2026-08-03, at v0.2.4 as of 2026-09-29. The video's frames show v0.1.3 and 1.5k stars; the count read 22,267 on 2026-09-29 and will keep moving.
- The video's format list is accurate. The README table lists Word (.doc, .docx, .docm), PowerPoint (.ppt, .pps, .pot, .pptx, .pptm, .ppsx, .ppsm), Excel (.xls, .xlsx, .xlsm, .xlsb), OpenDocument (.odt, .ods, .odp), RTF, EPUB, CSV and PDF.
- The one-model claim excludes PDF.
to_markdown_bytesreturns fromformats::pdf::to_markdownbefore the model exists, andto_documentis documented as unsupported for PDF. Scope any "fix it once" reasoning to the office formats. - Byte-based detection excludes CSV. Pass CSV bytes without a format argument and you get
Unsupported, because plain text carries no signature to read. - Before v0.2.4, scanned PDF pages were dropped without an error. If you are pinned below that version, audit your PDF output for missing pages before you trust it.
- Run
anydoc --ocr hostedonly with the knowledge that the entire file goes to Firecrawl Parse, a hosted service, with no page selection. - The benchmark is vendor-run, LLM-judged against LibreOffice renders of the first six pages, on a corpus that is not in the repository. Use it to shortlist, then run your own documents before you commit.
Resources
- firecrawl/anydoc. The repository from the video, with the README containing the supported-format table, the benchmark and the architecture diagram.
- src/lib.rs. Contains the PDF short-circuit in
to_markdown_bytesand the doc comment markingto_documentunsupported for PDF. - src/formats/detect.rs. The content-based detection module, including the note that CSV is never detected and falls back to the extension.
- v0.2.4 release notes. Documents that earlier versions silently dropped scanned PDF pages and introduced the
NeedsOcrerror. - bench/README.md. The benchmark harness, including how the LLM judge and the LibreOffice ground truth work.
- firecrawl/pdf-inspector. The separate Rust library anydoc hands every PDF to, which is why PDF output is governed by different code than the office formats.
- anydoc browser demo. Runs the library as WebAssembly so you can test your own document without installing anything or uploading it.
- crates.io/crates/anydoc. The Rust crate, at 0.2.4, useful for checking the current version against whatever you have pinned.
- firecrawl-anydoc on PyPI. The Python binding, which is the package name to install rather than
anydoc, even though the import isimport anydoc. - microsoft/markitdown. The breadth-first Python alternative, covering more file types including audio and images at lower fidelity per format.
- docling-project/docling. The ML-based alternative that reconstructs PDF layout and table structure, which is the right tool where anydoc's no-OCR, no-models position runs out.
- Unstructured-IO/unstructured. Apache-2.0 document ETL for LLM pipelines, and the closest competitor by format coverage in anydoc's benchmark.
- jgm/pandoc. The long-standing universal markup converter, GPL-2.0, whose 5-of-14 score in this benchmark reflects a different scope rather than a defect.
Published August 5, 2026. Writeup generated from a favorited TikTok.