> Source: https://txtfetch.com/blog/ocr-scanned-documents-api > Plain-text twin — every page on txtfetch.com has one. https://txtfetch.com/text --- # OCR scanned documents and images through one API call Feb 3, 2026 · 3 min read · ocr, api, llm-ingestion A large share of “documents” in any real intake pipeline aren’t documents in the structured sense at all. They’re images. A scanned invoice from a supplier who still uses a fax-to-email service. A photographed receipt from an expense app. A signed contract that was printed, signed, and re-scanned because someone didn’t trust the e-signature. All of them arrive as PNGs, JPEGs, or image-only PDFs with zero extractable text. If your extraction pipeline is built around a text-layer library (`pdftotext`, `python-docx`, anything that reads structured content), these files return nothing. Not an error, often just an empty string, which is the worst possible failure mode because it looks like success. ## Why OCR needs to be a first-class path, not a special case The usual fix is bolting Tesseract onto the side of a pipeline. First, detect that a file is an image. Then route it to a separate OCR step and merge the result back in. That’s a second dependency to install. It’s also a second set of language packs and model files to keep updated. And it’s a second code path that inevitably drifts out of sync with whatever error handling and retry logic the “real” text extraction path has. It also pushes a classification problem onto you: is this PDF text-based or scanned? Some are both: a scanned cover page stapled to a native PDF body. Get the routing wrong and you waste a Tesseract call on a document that didn’t need it. Or you silently skip OCR on a page that did. If you want to check which case one of your own files is, use [the free PDF text-layer checker](https://txtfetch.com/tools/pdf-text-check) before writing any routing logic. It answers that on the spot, in your browser, with no upload involved. ## One endpoint, format detection included txtfetch runs Tesseract as a Lambda layer alongside Apache Tika. It decides which one a request needs by inspecting the actual bytes, not the file extension or a flag you have to set. A `.jpg` receipt, a scanned `.pdf`, and a native `.docx` all go through the same call: ```bash curl -X POST https://api.txtfetch.com/v1/extract \ -H "Authorization: Bearer $TXTFETCH_KEY" \ -F file=@invoice-scan.jpg ``` Or, if the image is already sitting on a URL, say an attachment in an inbound email pipeline: ```bash curl -X POST "https://api.txtfetch.com/v1/extract?url=https://example.com/receipts/inv-4471.png" \ -H "Authorization: Bearer $TXTFETCH_KEY" ``` The response shape doesn’t change based on whether OCR ran or the file had a native text layer: ```json { "status": "success", "extracted_text": "INVOICE #4471\nBill to: Acme Corp\nTotal due: $1,240.00" } ``` Your downstream code (the code that reads `extracted_text` and does something with it) never needs to know which path a given file took. ## Handling OCR failures explicitly OCR is probabilistic in a way native text extraction isn’t: a blurry photo or a low-contrast fax can legitimately fail to yield usable text. txtfetch surfaces that as an explicit error rather than an empty success: ```json { "status": "error", "error": "No text could be extracted" } ``` That distinction matters in a pipeline. An empty string that reports `"status": "success"` gets embedded and indexed like any other chunk. That’s a silent gap in your retrieval. Nobody notices until a user asks about a document that’s technically “in” the system. An explicit error is something you can catch, log, and route to a re-scan queue or a human reviewer. ## Where this fits Scanned intake is rarely the whole pipeline. It’s the long tail: the supplier who still faxes, the older records that predate digital originals, the field team uploading phone photos of paperwork. Routing that long tail through the same `/v1/extract` call as everything else means one retry policy, one error format, and one thing to monitor. The alternative is a special-cased OCR branch that only gets tested when someone notices it’s broken. Have a batch of scanned documents you want to test against? [Get in touch](https://txtfetch.com/contact) and we’ll set you up with an API key. keep reading - **[Ingesting large documents and big batches without falling over](https://txtfetch.com/blog/batch-and-large-document-ingestion)** A 500-page PDF and a ten-thousand-file backfill stress the same two things: single-request time and concurrency. Jun 23, 2026 · 3 min read - **[Parsing DOCX, PPTX, and XLSX into clean text for LLM ingestion](https://txtfetch.com/blog/parse-office-docs-docx-pptx-xlsx-for-llms)** python-docx, python-pptx, and openpyxl each have their own API and their own way of failing on real files. One endpoint replaces all three. Mar 10, 2026 · 3 min read - **[Getting text out of video and audio: extraction vs. transcription](https://txtfetch.com/blog/text-from-video-and-audio)** Text in a media file can live in a sidecar, an embedded track, a container tag, or nowhere at all. Extraction and speech-to-text are different jobs. Aug 15, 2026 · 4 min read See also: [OCR for scanned documents & images →](https://txtfetch.com/extract/image) ## Try it on your own file. The free reader runs in your browser. Nothing gets uploaded. [Open the file reader →](https://txtfetch.com/tools/file-to-text) [Get an API key →](https://app.txtfetch.com/signup)