Home/Rebuild & Diagnose/Opening errors

There was an error opening this document

The same sentence appears for several quite different problems. Which one you have decides whether the file is recoverable, so the first useful step isn’t repairing it. It’s finding out.

Drop the PDF you want rebuilt

or choose a file

The qpdf engine (about 430 KB) is fetched only when you press the button, and it runs in this tab.

What the message is reporting

The reader reached the file and could not make sense of what it found. A PDF is read from both ends. The header at the top declares what the file is, and a marker near the bottom says where the index lives. The message appears when one of those two, or an object in between, fails to check out.

One sentence covers several faults, which is why the same wording can describe a download that stopped at 90 per cent and a deeply mangled file. The message is identical. The prospects are not.

The causes that produce it most often

  • An incomplete download, by a wide margin the most common. The transfer stopped early, so the reader holds a file whose end is missing.
  • Something that isn’t a PDF at all: a login page or an error page saved with a .pdf name, which happens when a download link sat behind a session that had already expired.
  • A zero byte or near empty file, from a failed export or a disk that filled up on the machine that produced it.
  • A file passed through software that mangled it: an upload that re-encoded it, a mail gateway that rewrote the attachment, an editor that crashed partway through a save.
  • Real structural damage inside an otherwise complete file, such as a damaged object or a stream whose declared length disagrees with its contents.

Four checks before you try to repair anything

  1. Download it again from the original source and compare the size. A second copy that is larger means the first one was truncated, which is the common case and the one where repair was never the answer.
  2. Open it in a different reader. Two readers disagreeing narrows the fault to the file rather than to the program.
  3. Look at the first four characters. A PDF begins with a percent sign followed by the letters P, D and F. If it begins with an angled bracket or a curly brace instead, you have an HTML or JSON page wearing a .pdf extension.
  4. Check the size against what you expect. A 300 page report that arrives at 41 KB did not arrive.

Then let the tool tell you which kind it is

Drop the file in and press the button. The tool checks the file as it arrived, rebuilds it, and checks the result. If the file is parseable you get a repaired copy and a note about what was found. If it’s not, you get a plain statement that the structure could not be read, and no output file at all.

That second outcome is the useful one. A tool that returns a file whatever happens has told you nothing. A tool that returns one only when the file was genuinely readable has told you the file is sound.

Rebuilds to no problems found
The file parsed and was rewritten cleanly. The export is the file to keep.
Rebuilds with warnings
Repaired, but a structure had to be worked around. Worth a re-export from source if it repeats.
Reports errors
The checks found faults. Compare before and after to see whether the rebuild improved anything.
Cannot read the structure
No output. The file is in the category no parser recovers, and the page says so rather than guessing.

When the file cannot be recovered

If the check reports that the structure could not be read, the damage is in the class no parser recovers from: truncated files, and files that have lost the pointer to their own table of contents. Every repair tool works by parsing, which is exactly why none of them helps here.

Go back to the source instead. Ask for a fresh copy, download again from the original link while signed in, or export once more from the document that produced it. That is the only route that ends in a correct file rather than a partly readable one.

A password protected file is a separate case. If the document was encrypted before it broke, repair and decryption are two jobs, and a file damaged in the way described here may have lost the dictionary either one needs. Try opening it with the password first. If a reader can open it at all, the structure survived and a rebuild is worth running.

Straight answers

Does this message always mean the file is damaged?

No. It’s what a reader shows when it can’t parse the file, and that includes files that were never PDFs. An HTML error page saved with a .pdf extension produces the identical sentence.

Can I fix it by renaming the file?

Renaming fixes the extension, not the contents. If the file really is a PDF whose download stopped early, a fresh download is the fix. If it genuinely is an HTML page, no rename changes what the bytes are.

Will rebuilding make the text selectable again?

If the text layer survived, it was never lost. A rebuild copies the content through without re-encoding it. If the file came back as images, that happened earlier, when it was scanned or exported, and rebuilding won’t undo it.

The file opens on my phone but not on my computer. Which one is right?

Both, usually. Phone readers are more forgiving of structural faults, so a file that opens there and fails on a stricter desktop reader is the parseable kind, and a rebuild is exactly the right move.

More on rebuild & diagnose

The full tool, with every option: Rebuild & Diagnose.