Home/Extract Text/Scanned PDFs

A scanned PDF has no text layer — here’s how to read it anyway

You opened the file, the words are right there on screen, and the extractor returns nothing. The words aren’t in the file — they’re in a photograph of a page that happens to be inside a PDF. That is diagnosable in five seconds, and for printed English it is fixable without uploading anything.

Drop the PDF you want the text from

or choose a file

Reports, papers, e-books, invoices. Scanned pages have no text layer; tick "OCR scanned pages" below to read them.

A PDF is a container, and it can hold either

The format has two ways to put words on a page. It can draw them as text, a list of characters with positions and a font, or it can place an image made by a scanner or a camera. Both look identical on screen and behave completely differently underneath.

Text can be selected, searched, reflowed and extracted. An image can only be looked at. When a scanner puts a page into a PDF, what lands in the file is one big picture per page, and a picture has no characters in it, no matter how clearly you can read them.

Three checks, none of them longer than a second

  • Drag-select across a word. If nothing highlights, there’s no text under it.
  • Press Ctrl+F and search for a word you can plainly see. A scan finds nothing.
  • Zoom to 400%. Text stays crisp because it’s drawn at your zoom level; a scan goes soft and blocky because it’s being enlarged.
A file can be mixed. Plenty of real documents are both, such as a scanned cover sheet stapled to a born-digital report. The extractor pulls text from the pages that have it and reports the ones that don’t rather than failing the whole file, and the OCR option runs only on those pages — the typed ones keep their exact characters.

What happens when you tick OCR

Leave the option off and the extractor reads the text layer and nothing else: instant, exact, local. A scan comes back empty because there is nothing to read, and the result panel names those pages rather than handing you a document with silent holes in it.

Tick it and a recognition engine runs over the pages that have no text layer — only those pages. It looks at pixels and returns characters. The engine is about 7 MB and downloads the first time you use it, not when the page loads, so a visitor who never opens a scan pays nothing for it.

Born-digital PDF
Text comes out of the text layer. Layout is approximate, but the characters are exact.
Scanned PDF
Nothing without OCR. With it ticked, clean print comes out close; handwriting and faint faxes do not.
Mixed document
Typed pages come from the text layer. OCR touches only the image pages.
Right-to-left or CJK text
The text layer comes out, but the reading order of a complex line can be wrong. OCR is English-only, so a CJK scan stays empty.

What this OCR is good at, and what it isn’t

It reads printed English. A clean 300 dpi scan of a typed page, or a PDF exported from a picture of one, comes through well enough to search, copy and quote. Handwriting, faint faxes, skewed pages and anything in another language come back wrong or empty, and the result panel names the pages where it found nothing. Treat the output as a draft rather than a transcript and read it against the original, because a wrong digit in a reference number looks exactly like a right one.

It costs one download, once, and only if you ask for it: roughly 7 MB of recognition engine and English model, fetched the first time you tick the box. A visitor who never opens a scan pays nothing, and the document never leaves the tab. No server sees it, which is the entire reason this runs here instead of somewhere with a queue.

Retyping three lines is often the fastest route. For a short document, reading it and typing it takes a couple of minutes and can’t fail. OCR pays off on a long file you need searchable; on a handful of lines the download and the checking cost more than the typing.

If the scan is bad enough that a human struggles

OCR isn’t the only step that gets harder as quality drops. A skewed page, heavy speckling from a photocopier, or a low-contrast fax will defeat most engines, and the ones that claim to handle it do so by guessing at characters.

What helps more than any tool is a better source. If there’s a paper original, scanning it again at 300 dpi in grayscale beats every post-processing option. If it came from someone else, asking them for the original digital file costs one email.

Straight answers

Can it OCR the scan for me?

Yes — tick OCR scanned pages (English). Roughly 7 MB of recognition engine and English model download the first time you use it, and the pages are read inside your tab. Nothing is uploaded and no account is involved.

Does it work on a scan in Chinese or another language?

No. The option is English-only, and each extra language is another model of several megabytes on top of the one already there. A French or CJK scan comes back empty rather than being fed to an engine that would guess at it.

Is OCR worth it for a single page?

Usually not. It costs a download the first time, and a wrong reading in a contract or an invoice costs more than the typing it saved. For a handful of lines, retyping wins; OCR earns its place on a long file you need searchable.

The preview shows some words. Why is that?

Mixed files are common. Those words are on the pages that genuinely carry a text layer. The empty pages are the images, and the result panel names them rather than leaving you to guess.

Does it matter that the text comes out with odd line breaks?

Somewhat. A PDF has no paragraphs, only characters at coordinates. The extractor groups them into lines by position, which is close to right most of the time and occasionally puts a caption into the middle of a column. Read the output before you rely on it.

Is the PDF I opened modified?

No. Nothing is written back. The extractor only reads, and the download is a separate text or Markdown file.

More on extract text

The full tool, with every option: Extract Text.