The vocabulary of text extraction, defined plainly.

These are the terms developers search for before they know they need an extraction API. They include what a PDF text layer actually is, why mojibake happens, and what OCR does and doesn't do. Each one links to the page with the full story.

Every entry below is one honest sentence. The full explanation, examples, and what to do about it live on the term's own page. Looking for a symptom instead of a definition? See fixes for broken extracted text.

See the term in real output.

Drop a file into the free reader and watch it happen.

Open the file reader →

Get an API key →