What the message is reporting
The reader reached the file and could not make sense of what it found. A PDF is read from both ends. The header at the top declares what the file is, and a marker near the bottom says where the index lives. The message appears when one of those two, or an object in between, fails to check out.
One sentence covers several faults, which is why the same wording can describe a download that stopped at 90 per cent and a deeply mangled file. The message is identical. The prospects are not.
The causes that produce it most often
- An incomplete download, by a wide margin the most common. The transfer stopped early, so the reader holds a file whose end is missing.
- Something that isn’t a PDF at all: a login page or an error page saved with a .pdf name, which happens when a download link sat behind a session that had already expired.
- A zero byte or near empty file, from a failed export or a disk that filled up on the machine that produced it.
- A file passed through software that mangled it: an upload that re-encoded it, a mail gateway that rewrote the attachment, an editor that crashed partway through a save.
- Real structural damage inside an otherwise complete file, such as a damaged object or a stream whose declared length disagrees with its contents.
Four checks before you try to repair anything
- Download it again from the original source and compare the size. A second copy that is larger means the first one was truncated, which is the common case and the one where repair was never the answer.
- Open it in a different reader. Two readers disagreeing narrows the fault to the file rather than to the program.
- Look at the first four characters. A PDF begins with a percent sign followed by the letters P, D and F. If it begins with an angled bracket or a curly brace instead, you have an HTML or JSON page wearing a .pdf extension.
- Check the size against what you expect. A 300 page report that arrives at 41 KB did not arrive.
Then let the tool tell you which kind it is
Drop the file in and press the button. The tool checks the file as it arrived, rebuilds it, and checks the result. If the file is parseable you get a repaired copy and a note about what was found. If it’s not, you get a plain statement that the structure could not be read, and no output file at all.
That second outcome is the useful one. A tool that returns a file whatever happens has told you nothing. A tool that returns one only when the file was genuinely readable has told you the file is sound.
- Rebuilds to no problems found
- The file parsed and was rewritten cleanly. The export is the file to keep.
- Rebuilds with warnings
- Repaired, but a structure had to be worked around. Worth a re-export from source if it repeats.
- Reports errors
- The checks found faults. Compare before and after to see whether the rebuild improved anything.
- Cannot read the structure
- No output. The file is in the category no parser recovers, and the page says so rather than guessing.
When the file cannot be recovered
If the check reports that the structure could not be read, the damage is in the class no parser recovers from: truncated files, and files that have lost the pointer to their own table of contents. Every repair tool works by parsing, which is exactly why none of them helps here.
Go back to the source instead. Ask for a fresh copy, download again from the original link while signed in, or export once more from the document that produced it. That is the only route that ends in a correct file rather than a partly readable one.
