Home/How it works

How it works

"Runs in your browser" is easy to print and hard to trust. This page lays out the whole mechanism: which libraries do what, where each byte goes, and the checks you can run yourself, plus the limits, described as plainly as the capabilities.

The libraries, each with one job

None of this is a PDF engine written from scratch, and none of it is fetched from somewhere else at runtime. The libraries below are bundled with the site and served from this domain, which is why the tools keep working with the network switched off. Three of them are WebAssembly modules rather than plain JavaScript, and each of those loads on demand only when a file needs it.

pdf-lib: rewriting the structure
Reads a PDF's object tree so pages can be merged, reordered, rotated, split, or dropped, and metadata edited out. This is the lossless route: nothing about a page's contents is reinterpreted, so text stays text and fillable fields stay fillable.
pdf.js (Mozilla): rendering pages
Draws a page onto a canvas to produce pixels. Everything that needs to see the picture rather than the data goes through it: the side-by-side comparison, burning a black bar into the page image so a redaction is real, and re-encoding scans for compression.
JSZip: packing results
When a tool produces many files at once, say a PNG per page, this bundles them into a single archive for download.
qpdf (compiled to WebAssembly): decryption fallback, rebuild, and encryption
A C++ PDF toolkit compiled to a .wasm module. It isn't used for ordinary files; it's fetched on demand only when the built-in handler can't read a file's encryption structure, when a file's structure is being rebuilt, or when you encrypt a file yourself. It still runs entirely in your tab, and the file never leaves the browser.
MozJPEG (compiled to WebAssembly): re-encoding scans
The JPEG encoder behind the compressor's image route. It only comes into play once pages are being redrawn as pictures, and there it beats the browser's own encoder: a page full of photographic detail is where it earns its keep.
tesseract (compiled to WebAssembly): reading a scan
The recognition engine behind the text extractor's “OCR scanned pages” option. It's the largest download on the site — roughly 7 MB, fetched only when you switch that option on — and it reads printed English. Handwriting and poor scans come back imperfect, which is why the tool warns you before it runs.

That's where the “pure JavaScript” line this page used to take ended, and it's worth saying why. Not because JavaScript can't do the job — it does most of it here. It's that all three Wasm modules are mature C and C++ projects, each built for exactly the awkward cases a from-scratch implementation handles worst: damaged structures, unusual encodings, scans that were never meant to be read by a machine. Compiling them puts already-tested code on that work, instead of asking you to trust a brand-new parser with the file you least want to lose.

One file, start to finish

  1. You choose a file. The browser hands the page a handle to a file on your disk. Nothing has moved yet; the page can't even see the bytes without asking.
  2. The page reads those bytes into memory in your tab. This is the step that replaces an upload on a normal site: same local operation, no network involved.
  3. The relevant library parses the document there. Any error about a malformed or encrypted file comes from your own browser reading your own file.
  4. Your operation runs. Either the structure route rewrites the object tree, or, if it needs the picture, pages are drawn to canvases and rebuilt from those images. Both outcomes are described below, including what each one costs.
  5. The result becomes a Blob in that same tab and is offered as a download from a local object URL. No server is waiting for it, so there's no copy anywhere else and nothing to recover later.

Check it yourself

  1. Unplug. Load any tool page, switch off Wi-Fi, then run a file through it. Every operation should complete, because the dependencies are already on your machine by then.
  2. Watch the network panel. Open DevTools, switch to Network, tick "Preserve log", drop a file in, and work through the tool. Every request you see resolves to this domain. A request to any other domain would be the thing this site promises not to do.
  3. Read what shipped. Open /sw.js. It's served like any other file and contains the complete list of everything the site caches for offline use: site assets only, pages, stylesheets, scripts. Auditing it takes no imagination, because the list is the list.

The parts that cost you something

Read this before you use a tool for something that has to survive scrutiny. These are trade-offs built into doing PDF work inside a browser, and each one shows up in the tool itself too.

Structure edits lose nothing; rasterized pages lose their text layer
Drawing pages as images is what makes a redaction permanent and lets a scan be compressed hard. It also means those pages stop being searchable and selectable, and that fillable form fields, comments, links, and bookmarks don't survive, because there's nothing left in the file to hold them.
Compression runs two routes and keeps the smaller result
The lossless route and the image route are both attempted, and you get whichever comes out smaller. On a fixed set of sample documents, the difference wasn't subtle: an image-heavy page with form fields shrank by roughly 95% and lost its eight form fields, while a twenty-page text-only sample grew by a factor of 461 on the image route and was never offered. Numbers that large are why the tool shows you both instead of picking quietly. Where keeping fields matters, there's a control to skip that route entirely.
Very large pages get quietly scaled down
Above roughly 24 megapixels per page, rasterizing reduces the resolution instead of failing. That keeps the browser alive, but an enormous page comes back smaller than it went in: fine for reading, wrong as an archival master.
OCR is one tool, in one language
Everywhere except the text extractor, the tools read text that's already in the file: a scan nobody has recognized has none, so extraction returns nothing and searching finds nothing. The extractor's “OCR scanned pages” option does recognize glyphs, but only printed English, and only as well as the scan allows — handwriting and poor scans come back wrong. Treat that output as a first pass you check, not a transcript you send.
Encryption stays encryption
Password-protected files open with the password you already have. Nothing here recovers one; that boundary is in the format, not a shortcut taken here.
The ceiling is your own machine
Everything is one tab doing the work, limited by the memory that tab can get. Work yields to the interface regularly so it never looks frozen, but thousands of pages at once is asking a lot; splitting usually gets there.

What is stored on your device

Almost nothing, but not literally nothing, so it's worth being precise: a light/dark theme preference, offline copies of the site's own files, and, when you start on one page and continue into a tool, the file's bytes handed over through local storage for at most ten minutes. Everything stays on your machine and disappears when you clear site data. The privacy policy names each one so you can go look.

Where each tool lands on all this

Merge PDFs

Join several files into one, drag to reorder, and pull just the pages you want from each.

Organize Pages

A thumbnail grid you can reorder, rotate, delete and duplicate, then export as a new file.

Compress

Lossless restructuring and image re-encoding, run separately. You keep whichever came out smaller.

Rebuild & Diagnose

A qpdf check first, then the internals get rewritten: cross-reference table rebuilt, unreferenced objects dropped, streams recompressed, and one last check. If a file is beyond saving, it tells you straight.

Redact

Box off account numbers, ID numbers, figures. The text underneath is destroyed, not covered up.

PDF ↔ Images

Export pages as high-resolution PNG or JPG in a ZIP, or combine a stack of scans into a PDF.

Watermark & page numbers

Tile a watermark or place it in one spot, Chinese included. Add page numbers in your own format.

Strip Metadata

See which author name, which machine, which software version a file is carrying. Then wipe it.

Crop PDF

Trim the white border off a scan, or let it snap to the content edges. Writes the page boundary, not the content.

Remove Password

Use the password you already have and get back a file you can print and merge. RC4 and AES-128/256. No cracking.

Encrypt PDF

Put a lock on the file: AES-256 encryption, separate passwords for opening and for permissions, and restrictions on printing, copying, form filling, and changes.

Sign & Stamp

Draw a signature, upload a stamp, or just type a name. Drag it where it goes and pull the corners.

Compare

Two versions, page by page. A pixel diff shows what moved; a line-level text diff says what the change was.

Light Edit

Cover a stale line, retype the right one, drop in an image. And an honest note about what that does and doesn't hide.

Extract Text

Pull the text layer out page by page into TXT or Markdown. Chinese included. Scanned pages can go through OCR — English only, and it tells you what it could not read.

Fill Forms

Read the fillable fields a PDF already has, fill them in, and export — either kept editable or flattened in place.

Sanitize PDF

Clean out the parts of a PDF that can run or reach out: JavaScript, launch actions, embedded files, external links, and attachments.

Extract Images

Export the images stored in a PDF rather than a rendering of the page. JPEG comes out byte-for-byte; other formats are re-encoded as PNG, all in a ZIP.

Batch

Drop in a whole folder, pick one action, strip metadata, sanitize, compress, repair, encrypt, unlock, or export text, embedded images, or page images, and get everything back in a ZIP.

Impose & Insert

Fit several pages onto one sheet (2-up / 4-up), arrange a booklet, or insert pages from another PDF — all aimed at printing and binding.

Still unsure whether a particular job behaves the way you need? Describe what you're trying to do and ask. There's one email address, and it's read by whoever wrote this.