Why the clipboard scrambles
A PDF has no paragraphs, no columns and no tables. It has text runs parked at coordinates, in whatever order the program that produced the file wrote them.
Copy-paste follows that order. On a single-column document it looks fine. On a two-column paper the runs alternate between columns, because that’s the sequence they were drawn in, and what lands in your editor is two articles shuffled into one.
What a tool can do about it, and what it can’t
The fix is to stop trusting the stored order and rebuild the order from the coordinates: group the runs that share a baseline, sort each group left to right, then step down the page.
- Page range: take the three pages you need instead of the whole report.
- Page boundaries: a marker between pages, so you can see where the layout changed.
- Repeated headers and footers: the filter looks only at the first three and last three lines of each page, and drops a line only when it repeats on at least 60% of them. A running head that also appears once as body text stays put.
- Heading guess: compares each line against the median body size and promotes what’s clearly larger. Two levels only, and it will occasionally promote a bold caption. The result panel says how many it promoted.
- Editable preview: the output is a text area on purpose. Fix the two lines the guess got wrong before you copy.
Rotated pages
Scans and exports often store a landscape page with a rotation flag instead of rotating the content. Extraction has to account for that, or the text comes out along the wrong axis, with two columns read as one improbable line. Both the reading order and the page-boundary detection follow the displayed orientation.
When one page is a scan and the rest aren’t
Mixed documents are common: a report with a scanned appendix, a contract with a photographed signature page. Extraction works page by page, so the typed pages come out and the scanned ones produce nothing.
An empty page in the output isn’t a failure. It means there was no text on it, only a picture of text. If you can select words on screen and they still don’t appear in the export, something else is going on and it’s worth reporting.
Getting it into the shape you need
Plain text when the destination is a note, a ticket, or a cell you’ll clean up by hand. Markdown when the destination understands structure: the export writes the file name as its title, one section per page, and the headings it inferred.
Bold, italics, lists, tables and links aren’t reconstructed. The page doesn’t say reliably which is which, and inventing them produces a file that looks structured and isn’t.
