Home/Extract Text/Copying text

Copy text out of a PDF

Ctrl+A, Ctrl+C, paste. That’s the entire operation when the file has a text layer. People end up at an online converter because the paste comes out interleaved, and a converter is the wrong answer to a layout problem.

Drop the PDF you want the text from

or choose a file

Reports, papers, e-books, invoices. Scanned pages have no text layer; tick "OCR scanned pages" below to read them.

Why the clipboard scrambles

A PDF has no paragraphs, no columns and no tables. It has text runs parked at coordinates, in whatever order the program that produced the file wrote them.

Copy-paste follows that order. On a single-column document it looks fine. On a two-column paper the runs alternate between columns, because that’s the sequence they were drawn in, and what lands in your editor is two articles shuffled into one.

What a tool can do about it, and what it can’t

The fix is to stop trusting the stored order and rebuild the order from the coordinates: group the runs that share a baseline, sort each group left to right, then step down the page.

  • Page range: take the three pages you need instead of the whole report.
  • Page boundaries: a marker between pages, so you can see where the layout changed.
  • Repeated headers and footers: the filter looks only at the first three and last three lines of each page, and drops a line only when it repeats on at least 60% of them. A running head that also appears once as body text stays put.
  • Heading guess: compares each line against the median body size and promotes what’s clearly larger. Two levels only, and it will occasionally promote a bold caption. The result panel says how many it promoted.
  • Editable preview: the output is a text area on purpose. Fix the two lines the guess got wrong before you copy.
Nothing is stitched into paragraphs, deliberately. Merging lines into flowing paragraphs means guessing where sentences end, and a wrong guess swallows the heading above it into the paragraph. An honest line break is easier to fix than a paragraph that has eaten its own title.

Rotated pages

Scans and exports often store a landscape page with a rotation flag instead of rotating the content. Extraction has to account for that, or the text comes out along the wrong axis, with two columns read as one improbable line. Both the reading order and the page-boundary detection follow the displayed orientation.

When one page is a scan and the rest aren’t

Mixed documents are common: a report with a scanned appendix, a contract with a photographed signature page. Extraction works page by page, so the typed pages come out and the scanned ones produce nothing.

An empty page in the output isn’t a failure. It means there was no text on it, only a picture of text. If you can select words on screen and they still don’t appear in the export, something else is going on and it’s worth reporting.

Getting it into the shape you need

Plain text when the destination is a note, a ticket, or a cell you’ll clean up by hand. Markdown when the destination understands structure: the export writes the file name as its title, one section per page, and the headings it inferred.

Bold, italics, lists, tables and links aren’t reconstructed. The page doesn’t say reliably which is which, and inventing them produces a file that looks structured and isn’t.

Straight answers

Can I just select the text on the page and copy it?

Yes, when it selects cleanly. This is for when it doesn’t: two-column layouts, tables, anything where the paste arrives shuffled.

Why is the output one line per line instead of proper paragraphs?

Because the tool refuses to guess where sentences end. Every line it sees is written on its own line, which is what you want when you’re about to check it.

Does it handle Chinese or Japanese?

Yes. Order is read from coordinates rather than from spaces or punctuation, so scripts that don’t put spaces between words work the same way.

What about a PDF with no text at all?

Then there’s no text layer to read, and the extractor will say so rather than guess. It can still help: tick the OCR option and a recognition engine runs over those pages inside your tab. It reads printed English, and it names the pages it couldn’t read.

More on extract text

The full tool, with every option: Extract Text.