> Source: https://txtfetch.com/diff > Plain-text twin — every page on txtfetch.com has one. https://txtfetch.com/text --- # The numbers are on /benchmarks. Here's the actual text. $ diff raw.txt txtfetch.txt — ten documents, the raw parser output on the left, the corrected text on the right, every difference marked. Nothing hidden. before you read the rest of this page This corpus is small and synthetic: ten recorded fixtures standing in for real Tika and VLM responses. What follows demonstrates the shape of each failure and its fix, not a claim of population-level accuracy. Every vendor in this market publishes a self-scored benchmark number; we'd rather show you the actual text and let you judge it yourself. Want evidence on a document of your own instead of taking our word for any of this? [Run it in the playground](https://txtfetch.com/playground) or [drop a file into the free converter](https://txtfetch.com/tools/file-to-text). Both read your file in your browser, and nothing is uploaded. the three hard ones ## Where Tika's structural recovery is weakest. A two-column PDF, a scanned invoice, and a wide pricing table make up this part of the corpus. These three carry a recorded premium-tier (vision-language model) response. The standard tier's gap on each one is wide enough to be worth showing. Q3 2026 Regional Performance Report (2-column PDF) `q3-2026-regional-performance.pdf` · application/pdf · left pane read via Apache Tika, **right pane at premium (VLM)** raw parser output ``` Q3 2026 Regional Performance Report Q3 was a strong quarter across every region, driven Support tickets rose 8% quarter over quarter, concentrated by the launch of the structured-output beta and continued in the API-key rotation flow; a fix shipped mid-quarter expansion in EMEA. Renewal rates held above 92% despite reduced volume by half within two weeks. the price increase that took effect in July. Region Revenue Growth North America $5.1M 11% EMEA $3.4M 24% APAC $1.8M 15% ``` what txtfetch returns ``` # Q3 2026 Regional Performance Report Q3 was a strong quarter across every region, driven by the launch of the structured-output beta and continued expansion in EMEA. Renewal rates held above 92% despite the price increase that took effect in July. Support tickets rose 8% quarter over quarter, concentrated in the API-key rotation flow; a fix shipped mid-quarter reduced volume by half within two weeks. | Region | Revenue | Growth | | --- | --- | --- | | North America | $5.1M | 11% | | EMEA | $3.4M | 24% | | APAC | $1.8M | 15% | ``` Tika's PDF parser reads the page in raw position order, not by column. Line one of the left column lands next to line one of the right, so two unrelated sentences run together mid-thought. Every word from the page survives. Only the order doesn't. The table in the right column degrades into one flattened line with no structure at all. The premium tier reads the page the way a person would instead of inferring order from coordinates, and recovers both paragraphs and the table. Reading order, standard vs. premium: 50.0% → 100.0%. See [/benchmarks](https://txtfetch.com/benchmarks) for the full methodology. [Columns out of order →](https://txtfetch.com/fixes/columns-out-of-order) Scanned Invoice #9042 (image-only PDF) `invoice-9042-scan.pdf` · application/pdf · left pane read via Apache Tika, **right pane at premium (VLM)** raw parser output ``` lnvoice #9O42 Bill to: Northwind Tradcrs, 44 Harbor R0w ltem Qty Unit Pricc Total Onsite installatlon 2 $45O.OO $9OO.OO Annual support pian 1 $1,2OO.OO $1,2OO.OO Tota1 due: $2,1OO.OO ``` what txtfetch returns ``` # Invoice #9042 Bill to: Northwind Traders, 44 Harbor Row | Item | Qty | Unit Price | Total | | --- | --- | --- | --- | | Onsite installation | 2 | $450.00 | $900.00 | | Annual support plan | 1 | $1,200.00 | $1,200.00 | Total due: $2,100.00 ``` This scan has no text layer, so the standard tier's automatic Tesseract OCR escalation is what reads it at all. Tesseract reads most of the page, but confuses look-alike characters (O for 0, l and I for 1) throughout. It also has no table model, so the two-line item table survives only as one run-on blur. The premium tier reads the same image through a vision-language model and gets both the characters and the table right. See /fixes/pdf-extracts-no-text for the OCR-escalation path this document takes. Table cell-F1 and reading order, standard vs. premium: 0.0% → 100.0%. See [/benchmarks](https://txtfetch.com/benchmarks) for the full methodology. [PDF extracts no text (the OCR escalation path) →](https://txtfetch.com/fixes/pdf-extracts-no-text) Q3 Product Catalog (5-column pricing table) `q3-product-catalog.pdf` · application/pdf · left pane read via Apache Tika, **right pane at premium (VLM)** raw parser output ``` Q3 Product Catalog Pricing reflects the July rate card; stock levels are warehouse totals as of quarter-end. Product Monthly Annual Stock WD-100 Widget Pro $29 $290 120 WD-200 Widget Mini $9 $90 340 WD-300 Widget Max $59 $590 45 ``` what txtfetch returns ``` # Q3 Product Catalog Pricing reflects the July rate card; stock levels are warehouse totals as of quarter-end. | SKU | Product | Monthly | Annual | Stock | | --- | --- | --- | --- | --- | | WD-100 | Widget Pro | $29 | $290 | 120 | | WD-200 | Widget Mini | $9 | $90 | 340 | | WD-300 | Widget Max | $59 | $590 | 45 | | WD-400 | Widget Nano | $5 | $50 | 610 | ``` Unlike the financial report above, Tika does detect a table here. It just recovers it wrong. The SKU column folds into the product name, so five columns become four. The last row falls off the page-boundary extraction entirely, with no error and no warning. One product goes silently missing from whatever index this feeds. The premium tier keeps all five columns and all four rows. Table cell-F1, standard vs. premium: 0.0% → 100.0%. See [/benchmarks](https://txtfetch.com/benchmarks) for the full methodology. [Tables losing structure →](https://txtfetch.com/fixes/tables-lose-structure) the other seven ## Office, HTML, and email — already close to exact. No recorded premium response exists for these seven. There's little for one to fix: a real structural model, or a parser that already reads the format correctly. Worth showing anyway. This is the other half of "delete the parser zoo," with receipts. One note on precision: the right pane for these seven is the human-checked ideal from `expected.json`. It is not a recording of what the engine returned. How close the standard tier actually lands on each is measured separately. It's 100.0% text similarity on six of the seven categories, and 92.3% on the DOCX. Those numbers, and the metric behind them, are on [/benchmarks](https://txtfetch.com/benchmarks). **Show all 7 documents** Invoice Follow-up (EML) `invoice-followup.eml` · message/rfc822 · left pane read via Apache Tika, **right pane: the human-checked expected.json** raw parser output ``` From: Priya Shah To: Billing Subject: Re: Invoice #7734 Hi team, Following up on the invoice below — could you confirm receipt? Description Qty Amount Consulting hours 12 $1,800.00 Travel 1 $340.00 Invoice total: $2,140.00 Due date: 2026-08-15 Thanks, Priya ``` the human-checked ideal ``` From: Priya Shah To: Billing Subject: Re: Invoice #7734 Hi team, Following up on the invoice below — could you confirm receipt? Description Qty Amount Consulting hours 12 $1,800.00 Travel 1 $340.00 Invoice total: $2,140.00 Due date: 2026-08-15 Thanks, Priya ``` Tika's RFC822 parser reads a plain email like this one without much to prove. Headers, body, and the small table all land where you'd expect. Sprint Status Update (Outlook MSG) `sprint-status-update.msg` · application/vnd.ms-outlook · left pane read via Apache Tika, **right pane: the human-checked expected.json** raw parser output ``` From: Dana Lee To: Engineering Subject: Sprint Status Update Hi all, Here is where things stand heading into the demo on Friday. Benchmark harness: on track Corpus curation: on track CI wiring: at risk, needs a second reviewer Ping me if you want to pair on the CI job. Thanks, Dana ``` the human-checked ideal ``` From: Dana Lee To: Engineering Subject: Sprint Status Update Hi all, Here is where things stand heading into the demo on Friday. Benchmark harness: on track Corpus curation: on track CI wiring: at risk, needs a second reviewer Ping me if you want to pair on the CI job. Thanks, Dana ``` Legacy Outlook .msg goes through Tika's OutlookExtractor instead, and it's just as unremarkable here: headers, prose, a short list, nothing lost. txtfetch Pricing Page (HTML) `pricing.html` · text/html · left pane read via Apache Tika, **right pane: the human-checked expected.json** raw parser output ``` Pricing Turn any document into clean plain text over a simple HTTP API. pricing tiers diagram Plans Plan Price Quota Hobby $0 500 docs/mo Developer $19 10,000 docs/mo Scale $99 100,000 docs/mo Questions? Use the contact form & skip the 'support' email. ``` the human-checked ideal ``` Pricing Turn any document into clean plain text over a simple HTTP API. pricing tiers diagram Plans Plan Price Quota Hobby $0 500 docs/mo Developer $19 10,000 docs/mo Scale $99 100,000 docs/mo Questions? Use the contact form & skip the 'support' email. ``` Tika strips the inline script and style noise and keeps the pricing table. An ordinary HTML page like this one is already a solved problem. Q3 2026 Board Update (DOCX) `board-update-q3-2026.docx` · application/vnd.openxmlformats-officedocument.wordprocessingml.document · left pane read via Apache Tika, **right pane: the human-checked expected.json** raw parser output ``` Q3 2026 Board Update This update covers Q3 2026 performance across all product lines. Highlights Annual recurring revenue grew 14% quarter over quarter Shipped the structured-output beta Markdown mode Element JSON mode Support headcount increased to 6 engineers Revenue by Plan Plan MRR Growth Hobby $0 n/a Developer $18,200 11% Scale $41,900 19% Full detail is available in the appendix & supporting spreadsheets. ``` the human-checked ideal ``` Q3 2026 Board Update This update covers Q3 2026 performance across all product lines. Highlights Annual recurring revenue grew 14% quarter over quarter Shipped the structured-output beta Markdown mode Element JSON mode Support headcount increased to 6 engineers Revenue by Plan Plan MRR Growth Hobby $0 n/a Developer $18,200 11% Scale $41,900 19% Full detail is available in the appendix & supporting spreadsheets. ``` A .docx carries a real structural model, not pixels to guess at. Tika's OOXML parser has an easy time of it: the headings, the nested bullet list, and the table all come through. S6 Launch Deck (PPTX) `s6-launch-deck.pptx` · application/vnd.openxmlformats-officedocument.presentationml.presentation · left pane read via Apache Tika, **right pane: the human-checked expected.json** raw parser output ``` Benchmarking the Quality Climb txtfetch S6 — accuracy harness Why it matters Structured output needs proof, not just a demo VLM tier cost only pays off if quality actually improves What ships Committed corpus across 10 categories Reproducible CLI + committed report ``` the human-checked ideal ``` Benchmarking the Quality Climb txtfetch S6 — accuracy harness Why it matters Structured output needs proof, not just a demo VLM tier cost only pays off if quality actually improves What ships Committed corpus across 10 categories Reproducible CLI + committed report ``` Tika reads one slide per page, title and bullets in order. A clean deck like this one needs nothing more than the standard tier. FY2026 Budget Forecast (XLSX) `fy2026-budget-forecast.xlsx` · application/vnd.openxmlformats-officedocument.spreadsheetml.sheet · left pane read via Apache Tika, **right pane: the human-checked expected.json** raw parser output ``` Opex Line item Q1 Q2 Cloud hosting 4200 4600 Payroll 61000 64500 Summary Total 65200 69100 ``` the human-checked ideal ``` Opex Line item Q1 Q2 Cloud hosting 4200 4600 Payroll 61000 64500 Summary Total 65200 69100 ``` Two sheets, two real tables. Tika reads .xlsx's actual cell model rather than inferring rows and columns from visual spacing. Security Brief (single-column PDF) `security-brief.pdf` · application/pdf · left pane read via Apache Tika, **right pane: the human-checked expected.json** raw parser output ``` Security Brief This document describes the security posture of the txtfetch extraction pipeline. Every document is treated as hostile until parsed: zip-bomb guards, SSRF-guarded fetches, and a hard extraction budget all run before Tika ever sees the bytes. Threat Model Untrusted input arrives as arbitrary bytes from a URL fetch or a direct upload. Secrets and document content are never logged; only allowlisted metadata fields reach stdout. ``` the human-checked ideal ``` Security Brief This document describes the security posture of the txtfetch extraction pipeline. Every document is treated as hostile until parsed: zip-bomb guards, SSRF-guarded fetches, and a hard extraction budget all run before Tika ever sees the bytes. Threat Model Untrusted input arrives as arbitrary bytes from a URL fetch or a direct upload. Secrets and document content are never logged; only allowlisted metadata fields reach stdout. ``` A born-digital, single-column PDF with a real text layer. Reading order is already linear, so there's nothing here for a vision model to fix. how to read this Every document above is committed at `api/extract/bench/corpus//` in the txtfetch repo. It includes a recorded `tika.txt` (raw parser output), a human-checked `expected.json`, and, for the three hard PDFs, a recorded `vlm.md`. `make bench` re-runs the real extraction engine against the same fixtures and regenerates the committed report. It uses a dependency-injected mock Tika and mock VLM, with no JVM and no network calls. For the full methodology, every metric definition, and the per-category numbers these documents feed into, see [/benchmarks](https://txtfetch.com/benchmarks). For how the whole pipeline fits together, see [/how-it-works](https://txtfetch.com/how-it-works). faq **Is this corpus proof that txtfetch is this accurate on my documents?**: No. It's ten small, synthetic documents: recorded fixtures standing in for real Tika and VLM responses. They demonstrate the shape of a handful of known failure modes, and their fix. They don't claim a population-level accuracy figure. Run your own file through /playground or the free file-to-text tool if you want evidence on your documents specifically. **Why do the seven "easy" documents look almost identical on both sides?**: Because they mostly are. Office files, HTML, and email carry a real structural model. Tika's own OOXML/RFC822/HTML parser already reads them cleanly. There's very little left for a vision-language model to improve. The right pane for those seven is expected.json's human-checked ideal, not a recorded engine response. /benchmarks measures how close the standard tier actually lands: 100.0% text similarity on six of the seven categories, and 92.3% on the DOCX. That's the honest case for txtfetch: one endpoint that already gets these formats right, without you wiring up seven separate parsers to find that out. **Does the premium tier really score 100% on the three hard PDFs?**: On these three recorded documents, yes, because the recorded vlm.md response is the human-checked ideal recovery for each one, not a general accuracy ceiling. Treat it as "the VLM read this page correctly," not "the VLM is always perfect." See /benchmarks for the full per-category numbers and their caveats. **Where do these files actually live, and how do I regenerate the numbers?**: Every document is committed at api/extract/bench/corpus// in the txtfetch repo. It includes a recorded tika.txt (raw parser output), a human-checked expected.json, and, for the three hard PDFs, a recorded vlm.md. `make bench` re-runs the real extraction engine against the same fixtures and regenerates the committed report at api/extract/bench/REPORT.md. It uses a dependency-injected mock Tika and mock VLM, with no JVM and no network calls. ## Stop parsing. Start shipping. Create an account and get an API key in minutes. The free Hobby plan needs no card. [Get started →](https://app.txtfetch.com/signup)