<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>txtfetch changelog</title><description>A dated record of what txtfetch shipped, with a link to the page that proves each entry.</description><link>https://txtfetch.com</link><item><title>Extract from Drive, SharePoint, S3, and more</title><link>https://txtfetch.com/changelog/#2026-09-06-sources</link><guid isPermaLink="true">https://txtfetch.com/changelog/#2026-09-06-sources</guid><description>The /sources page names every place a document can live before txtfetch reads it. It states which ones hand over a URL, and which need a file upload.</description><pubDate>Sun, 06 Sep 2026 00:00:00 GMT</pubDate></item><item><title>A free tool to clean up damaged extracted text</title><link>https://txtfetch.com/changelog/#2026-09-05-clean-extracted-text</link><guid isPermaLink="true">https://txtfetch.com/changelog/#2026-09-05-clean-extracted-text</guid><description>Paste text that came out of any extractor and see the damage signals. Each one links to the fix that explains it, and nothing is uploaded.</description><pubDate>Sat, 05 Sep 2026 00:00:00 GMT</pubDate></item><item><title>What pgvector, Pinecone, Qdrant, Chroma, and Weaviate need next</title><link>https://txtfetch.com/changelog/#2026-09-04-ingest</link><guid isPermaLink="true">https://txtfetch.com/changelog/#2026-09-04-ingest</guid><description>txtfetch stops at extracted text. /ingest covers the chunk, embed, and upsert step for five vector stores, cited and dated.</description><pubDate>Fri, 04 Sep 2026 00:00:00 GMT</pubDate></item><item><title>How txtfetch compares to five open-source parsers</title><link>https://txtfetch.com/changelog/#2026-09-03-compare-libraries</link><guid isPermaLink="true">https://txtfetch.com/changelog/#2026-09-03-compare-libraries</guid><description>/compare/libraries checks txtfetch against Docling, MarkItDown, PyMuPDF4LLM, Marker, and MinerU on licence, GPU need, format coverage, and OCR.</description><pubDate>Thu, 03 Sep 2026 00:00:00 GMT</pubDate></item><item><title>Document-language pages, with the honest OCR limit stated</title><link>https://txtfetch.com/changelog/#2026-09-01-languages</link><guid isPermaLink="true">https://txtfetch.com/changelog/#2026-09-01-languages</guid><description>/languages covers nine document languages: legacy encodings, chunking risk, and where OCR support actually stops.</description><pubDate>Tue, 01 Sep 2026 00:00:00 GMT</pubDate></item><item><title>Stopped promising guaranteed reading order on multi-column PDFs</title><link>https://txtfetch.com/changelog/#2026-08-30-solutions-reading-order-fix</link><guid isPermaLink="true">https://txtfetch.com/changelog/#2026-08-30-solutions-reading-order-fix</guid><description>/solutions claimed txtfetch always kept multi-column PDFs in reading order. It does not: the engine gets most PDFs right, and /fixes/columns-out-of-order covers the rest. The page now says so.</description><pubDate>Sun, 30 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Wire txtfetch into n8n, Zapier, Make, Airflow, or S3 and Lambda</title><link>https://txtfetch.com/changelog/#2026-08-29-integrations</link><guid isPermaLink="true">https://txtfetch.com/changelog/#2026-08-29-integrations</guid><description>No plugin ships for any of them. /integrations shows the HTTP step each platform already has, pointed at txtfetch.</description><pubDate>Sat, 29 Aug 2026 00:00:00 GMT</pubDate></item><item><title>An honest answer to &apos;why not self-host Tika?&apos;</title><link>https://txtfetch.com/changelog/#2026-08-24-self-hosted-tika</link><guid isPermaLink="true">https://txtfetch.com/changelog/#2026-08-24-self-hosted-tika</guid><description>Apache Tika is free. /compare/self-hosted-tika shows the pinned versions, the from-source Tesseract build, and the runtime limits behind the question.</description><pubDate>Mon, 24 Aug 2026 00:00:00 GMT</pubDate></item><item><title>An extraction glossary</title><link>https://txtfetch.com/changelog/#2026-08-17-glossary</link><guid isPermaLink="true">https://txtfetch.com/changelog/#2026-08-17-glossary</guid><description>/glossary defines PDF text layers, OCR, mojibake, OOXML, and chunking in plain language. Each term links to its full page.</description><pubDate>Mon, 17 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Raw parser output next to txtfetch&apos;s output, word for word</title><link>https://txtfetch.com/changelog/#2026-08-16-diff</link><guid isPermaLink="true">https://txtfetch.com/changelog/#2026-08-16-diff</guid><description>/diff shows ten committed documents with the raw parser text next to the corrected text, with every difference marked.</description><pubDate>Sun, 16 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Audio and video get their own extraction page</title><link>https://txtfetch.com/changelog/#2026-08-15-extract-captions</link><guid isPermaLink="true">https://txtfetch.com/changelog/#2026-08-15-extract-captions</guid><description>/extract/captions states exactly what comes back from an .srt, .vtt, .mp4, .mp3, or .mkv file. A file with no caption track and no tag data returns an error, not an empty success.</description><pubDate>Sat, 15 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Corrected the &apos;1,000+ formats&apos; claim</title><link>https://txtfetch.com/changelog/#2026-08-13-formats-coverage</link><guid isPermaLink="true">https://txtfetch.com/changelog/#2026-08-13-formats-coverage</guid><description>That figure counted every media type Tika can detect, not every one it can parse. /formats/coverage now lists the real, checked count, measured against the exact jar this build ships.</description><pubDate>Thu, 13 Aug 2026 00:00:00 GMT</pubDate></item><item><title>One drop zone for any file type</title><link>https://txtfetch.com/changelog/#2026-08-12-file-to-text</link><guid isPermaLink="true">https://txtfetch.com/changelog/#2026-08-12-file-to-text</guid><description>/tools/file-to-text detects a file&apos;s real type from its bytes and extracts its text in the browser, with no upload.</description><pubDate>Wed, 12 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Free readers for legacy Word, Excel, and PowerPoint files</title><link>https://txtfetch.com/changelog/#2026-08-10-legacy-office-readers</link><guid isPermaLink="true">https://txtfetch.com/changelog/#2026-08-10-legacy-office-readers</guid><description>/tools/doc-to-text reads the 97-2003 binary Office formats (.doc, .xls, .ppt) in the browser, alongside the existing DOCX/XLSX/PPTX tools.</description><pubDate>Mon, 10 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Check whether a scan will OCR cleanly before you send it</title><link>https://txtfetch.com/changelog/#2026-08-07-image-ocr-check</link><guid isPermaLink="true">https://txtfetch.com/changelog/#2026-08-07-image-ocr-check</guid><description>/tools/image-ocr-check measures resolution, focus, skew, and inversion from a photo&apos;s actual pixels, in the browser.</description><pubDate>Fri, 07 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Stopped claiming the API returns media metadata it has no field for</title><link>https://txtfetch.com/changelog/#2026-08-06-subtitles-metadata-fix</link><guid isPermaLink="true">https://txtfetch.com/changelog/#2026-08-06-subtitles-metadata-fix</guid><description>Copy on /tools/subtitles-to-text said the API returns a title, artist, duration, and codec for media files. It has no such field. A success carries the extracted text, the content type, the byte and character counts, and whether OCR ran.</description><pubDate>Thu, 06 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Read .eml and Outlook .msg files, attachments included</title><link>https://txtfetch.com/changelog/#2026-08-04-email-to-text</link><guid isPermaLink="true">https://txtfetch.com/changelog/#2026-08-04-email-to-text</guid><description>/tools/email-to-text extracts a message and its attachments recursively, in the browser.</description><pubDate>Tue, 04 Aug 2026 00:00:00 GMT</pubDate></item><item><title>In-browser DOCX, XLSX, and PPTX readers</title><link>https://txtfetch.com/changelog/#2026-08-03-office-to-text-tools</link><guid isPermaLink="true">https://txtfetch.com/changelog/#2026-08-03-office-to-text-tools</guid><description>/tools/docx-to-text and its XLSX and PPTX siblings read modern Office files in the browser. No upload, no API call.</description><pubDate>Mon, 03 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Every page on this site is also plain text</title><link>https://txtfetch.com/changelog/#2026-08-01-plain-text-mirror</link><guid isPermaLink="true">https://txtfetch.com/changelog/#2026-08-01-plain-text-mirror</guid><description>Append .md to any URL for a plain-text twin. /llms.txt indexes the whole site by section, and /llms-full.txt is the whole corpus in one file.</description><pubDate>Sat, 01 Aug 2026 00:00:00 GMT</pubDate></item><item><title>A page for text that extracted but came out wrong</title><link>https://txtfetch.com/changelog/#2026-07-31-fixes</link><guid isPermaLink="true">https://txtfetch.com/changelog/#2026-07-31-fixes</guid><description>/fixes covers symptoms like mojibake, missing spaces, and scrambled columns: what causes each one, and how to fix it.</description><pubDate>Fri, 31 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Use-case pages: RAG ingestion, search indexing, document workflows</title><link>https://txtfetch.com/changelog/#2026-07-21-solutions</link><guid isPermaLink="true">https://txtfetch.com/changelog/#2026-07-21-solutions</guid><description>/solutions shows the same API applied to different jobs, with the parts of the response each job actually uses.</description><pubDate>Tue, 21 Jul 2026 00:00:00 GMT</pubDate></item><item><title>Developer docs: quickstarts, async jobs, idempotency, errors</title><link>https://txtfetch.com/changelog/#2026-07-19-docs</link><guid isPermaLink="true">https://txtfetch.com/changelog/#2026-07-19-docs</guid><description>/docs covers authentication and both ways to call the API. It also covers response formats, async jobs and webhooks, idempotency keys, and the full error reference.</description><pubDate>Sun, 19 Jul 2026 00:00:00 GMT</pubDate></item><item><title>A public status page</title><link>https://txtfetch.com/changelog/#2026-07-16-status</link><guid isPermaLink="true">https://txtfetch.com/changelog/#2026-07-16-status</guid><description>/status shows current status and 90-day uptime history for the extraction API. No login is required.</description><pubDate>Thu, 16 Jul 2026 00:00:00 GMT</pubDate></item><item><title>A dedicated page for every format</title><link>https://txtfetch.com/changelog/#2026-07-16-extract-format-pages</link><guid isPermaLink="true">https://txtfetch.com/changelog/#2026-07-16-extract-format-pages</guid><description>/extract lists format-specific pages covering PDF, Office, email, HTML, images, and more, each with its own working curl.</description><pubDate>Thu, 16 Jul 2026 00:00:00 GMT</pubDate></item></channel></rss>