Most PDF bugs never show up in a demo. A hand-picked test file, a sample invoice, a contract you wrote yourself: these render fine almost by definition, because they were made by tools that follow the rules. The bugs that matter show up on document #1,142, a PDF exported by a fifteen-year-old scanner, a spreadsheet printed by an unusual print driver, a scan that was rotated by hand in a program nobody remembers the name of. We test our PDF engine against thousands of real, messy documents for exactly this reason, and the failures it finds are rarely the ones you would have thought to write a test for.
Why we test against real documents, not synthetic ones
The PDF format is a specification, but in practice it is closer to a dialect continuum. Almost every PDF producer (office suites, scanners, print drivers, other PDF tools re-saving a file) takes some liberty with the spec, and viewers have spent decades quietly tolerating those liberties so documents keep opening. A PDF engine that only handles textbook-correct files will pass every test it writes for itself and still fail on a meaningful slice of what people actually upload.
So alongside unit tests, we run our engine across a large, growing corpus of real documents: specs, scanned forms, exported spreadsheets, presentations converted to PDF, government filings, and compare what comes out against what a reference reader would show. The goal is not to hit a percentage and stop. It is to keep finding the next category of document that behaves in a way our first thousand files never did.
The font that told us the wrong story
Here is a fact most people never need to know: a PDF does not trust the fonts it embeds to describe their own letter spacing. Every font carries its own internal measurements, but the PDF format also stores a separate, explicit list of how wide each character should be on the page, and by the specification, a viewer is supposed to trust that separate list over the font itself. It exists so a document can be laid out correctly without a viewer having to inspect the font program at all.
Most of the time nobody notices this distinction, because the two numbers agree. But across a large enough corpus, they disagree often enough to matter, usually from documents produced by office and spreadsheet software, whose exporters do not always write that width list in the way the rest of a viewer expects to find it. Get this wrong and the symptom is subtle: letters sit slightly too close or too far apart, columns in a table drift out of alignment, or two adjacent words silently merge into one. Nothing crashes. It just looks a little wrong, in a way that is easy to write off as a font issue rather than a layout one, until you check it against a hundred similar files and see the same drift every time.
“Encrypted” doesn't always mean what you think
PDF encryption is really two separate locks: a password to open the file, and a separate password governing what you're allowed to do once it's open, such as printing, copying text, or editing. A document can restrict all of those and still open for anyone, with no password prompt at all, because only the second lock was ever set. To a naive reader, both cases look identical: the file reports itself as encrypted either way.
Across our corpus, this turned out to be the dominant real-world case by a wide margin. The overwhelming majority of “encrypted” documents we tested needed no password at all. A tool that treats every encrypted-looking file as password-protected ends up prompting real users for a password that does not exist, on documents they should have been able to open in one click. The fix is not complicated once you know to look for it, but you only find it by testing against documents that actually carry this pattern, because it is not the kind of thing you'd think to construct by hand.
Pages that render upside down (and why some viewers are quietly guessing)
A PDF page can carry an explicit rotation instruction: display this page turned 90 or 180 degrees from how its contents are actually drawn. Scanned documents lean on this constantly. A page scanned sideways is stored sideways, with a flag telling every viewer to turn it before showing it to a human. Ignore that flag, or apply it in the wrong place in the rendering pipeline, and the whole page comes out sideways or mirrored, so readable content becomes a hidden implementation detail instead of something a person can actually read.
What makes this category interesting is how differently viewers disagree at the edges: nested rotations, a rotation applied to part of a page's content but not its annotations, or a scan whose underlying text layer is itself rotated independently of the page. These are exactly the cases where “it worked when I tested it” and “it works” turn out to be different claims, and where a large, varied corpus earns its keep, because these combinations show up in maybe one document in a few hundred, never in the first ten you'd think to try.
What breaks at 1,600 files that never breaks at 10
None of the categories above are exotic once you know about them. They are public facts about how the PDF format works, documented in the specification for anyone who reads it closely. What changes at scale isn't the existence of these cases; it is how often they combine, and how quickly the “it works on my test files” signal stops meaning anything. A handful of hand-picked documents will never surface a font whose declared spacing disagrees with itself, a file that is encrypted but not really locked, and a scan whose text runs upside down inside a rotated page, all at once, in the same document. A few thousand real ones will, sooner or later.
That is really the argument for testing this way at all. Every fix we make gets checked back against the whole corpus, not just the file that surfaced it, because a fix that only solves the document in front of you is a fix that will need to be re-solved the next time a similar-but-different file shows up. The corpus is the thing that keeps a fix from being a one-off.
The takeaway
A document engine that only has to survive a demo will always look finished. The honest measure is how it holds up against the documents nobody curated: the scans, the exports, the fifteen-year-old files that were never meant to be a test case for anything. That is the standard we hold our own in-browser PDF engine to, and it is why tools like unlocking a PDF or fixing a sideways scan are built to handle the document you actually have, not the one a spec would have preferred you upload.