PDF Compression Methods Explained: Lossy, Lossless, and Everything Between
Click "compress," and a few seconds later a smaller file appears. Most people never ask what happened in between, and for everyday use, that's completely fine. But if you've ever wondered why compressing a photo-heavy brochure behaves nothing like compressing a scanned contract, or why one compressor leaves your file looking identical while another leaves it visibly softer, the answer lives in which specific compression methods are being applied and how they're combined.
This is a closer, more technical look at PDF compression methods than you'll usually find — the actual mechanics of lossy versus lossless compression, what downsampling and re-encoding really do to an image, how font subsetting quietly saves space without you noticing, and how object and stream compression clean up the file's internal structure. We'll also get specific about which of these methods matters most depending on what kind of document you're dealing with.
Lossy vs. Lossless Compression: The Core Distinction
Every compression technique falls into one of two categories, and understanding the difference explains almost everything else about how PDF compression behaves.
Lossless compression shrinks data without discarding any of it. The original information can be perfectly reconstructed from the compressed version — nothing is thrown away, it's just represented more efficiently. This is the kind of compression applied to text, fonts, and the PDF's internal structure, which is exactly why compressing a text-heavy document rarely produces any visible change: the text itself isn't being altered, just packed more efficiently.
Lossy compression shrinks data by selectively discarding information that's judged unlikely to be noticeable, and it cannot be perfectly reversed — some detail is genuinely gone. This is the category most image compression falls into, since photographic data is large enough that meaningful size reduction usually requires giving up some fine detail in exchange for a much smaller file.
The practical implication: a document that's mostly text can be compressed almost entirely with lossless methods, meaning close to zero visible tradeoff. A document that's mostly images has to lean on lossy methods to see any meaningful size reduction, which means there's a real (if often invisible, when done well) tradeoff involved.
Most PDF compressors blend both approaches automatically, applying lossless techniques to the structural and text elements of the file while applying calibrated lossy compression to embedded images — which is exactly the approach behind NanPDF's Compress PDF tool, rather than treating the whole document as one undifferentiated blob of data.
Downsampling: Reducing Image Resolution
Downsampling is probably the single most impactful compression method for image-heavy or scanned PDFs, and it's conceptually simple: it reduces the number of pixels stored per inch of an image.
A photo embedded at 300 DPI (dots per inch) contains far more pixel data than a screen typically needs to display it clearly — most monitors and phone screens render comfortably at 72 to 150 DPI for on-screen viewing. Downsampling an image from 300 DPI down to 150 DPI, for instance, roughly cuts the pixel count to a quarter of the original, because resolution reduction applies in both directions of the image (width and height), not just one.
This is lossy by nature — once pixels are discarded, they're gone — but done within a sensible range, the visual difference is imperceptible on a screen, since you're only removing detail beyond what the display could show anyway. The risk zone is pushing downsampling too aggressively on content that will later be zoomed into or printed, where the missing pixel data becomes visible as softness or blur.
Downsampling has essentially no effect on print-ready files if applied carelessly — or rather, it has too much effect, which is why print-destined PDFs typically need lighter downsampling settings or none at all, preserving resolution well beyond what a screen requires.
Re-Encoding: Changing How Image Data Is Stored
Re-encoding is a separate method from downsampling, though the two are often applied together. Instead of reducing how many pixels an image has, re-encoding changes how efficiently those pixels are stored.
Different image compression algorithms make different tradeoffs between file size and quality retention for the same pixel dimensions. A photograph re-encoded with a more efficient algorithm can end up considerably smaller at the same resolution, with a quality loss so slight it's very difficult to detect by eye, especially at typical viewing sizes. Push the re-encoding strength too far, though, and you start to see the telltale signs of aggressive lossy compression — blocky artifacts around sharp edges, color banding in smooth gradients, or a slightly "smudged" look in fine detail.
This is why compression strength settings in a good tool matter: light, medium, and strong settings usually correspond to different re-encoding aggressiveness, letting you choose where on that spectrum a given document should sit. A scanned certificate you'll only ever view on screen can tolerate a fairly aggressive re-encoding setting. A product photo in a sales brochure probably can't, if visual polish matters to the document's purpose.
Font Subsetting: Only Keeping What's Actually Used
Font subsetting is a quieter but genuinely effective compression method, and it's entirely lossless — nothing about how the document looks changes at all.
A full font file contains every character, weight, and style variant the typeface offers, often including characters from multiple alphabets and symbol sets that most documents never use. When a PDF embeds a complete font file just to render, say, 200 words of body text, it's carrying an enormous amount of unused data along for the ride.
Subsetting solves this by scanning the document for exactly which characters actually appear, then embedding only those specific glyphs rather than the entire font. A document using a decorative headline font for just a title and a handful of section headers might only need fifty or sixty characters from that font, instead of the several thousand a complete font file would include.
The size savings from subsetting scale with how many custom or non-standard fonts a document uses. A single-font, plain-text PDF sees relatively little benefit, since there's not much font data to begin with. A design-heavy document using several distinct typefaces can see a meaningfully smaller file purely from subsetting, with zero visual difference, since every character that's actually displayed is still fully and correctly embedded.
Object and Stream Compression: Cleaning Up the File's Structure
Beneath the visible content, a PDF is built from a collection of internal objects — page descriptions, font references, image data, annotations, and metadata — each stored as a data stream within the file. Object and stream compression targets this underlying structure rather than the visible content itself.
Stream compression applies lossless data compression to these internal streams, similar in spirit to how a ZIP file shrinks data without altering it. This is nearly always safe to apply aggressively, since it doesn't touch the actual visual content at all — it just packs the underlying data more efficiently.
Object deduplication removes redundant copies of the same internal object. This matters most in documents that have been edited, merged, or exported multiple times, where the same font reference, color profile, or image might exist in duplicate within the file structure without ever being cleaned up. Removing these duplicates can meaningfully shrink a file with zero visible impact, particularly for PDFs that have been through several rounds of editing or combining.
Metadata and unused object stripping clears out orphaned data — old revision history, unused thumbnails, embedded comments or annotations left over from editing — that adds file weight without contributing anything to what the reader actually sees.
Together, these structural methods are almost entirely lossless and are a big part of why a compressor can sometimes shrink a file noticeably even when the visible images and text look completely untouched.
Which Method Matters Most for Which Document Type
Not every compression method carries equal weight for every kind of file. Matching the method to the document is where meaningful size reduction actually comes from.
Photo-heavy brochures and marketing materials. Downsampling and re-encoding do almost all the work here, since the bulk of the file size lives in embedded images. Font subsetting adds a smaller but real benefit if the design uses multiple branded typefaces. Because these documents are often visual by purpose, it's worth using a moderate rather than aggressive compression setting to preserve polish.
Text-heavy contracts and reports. Object and stream compression, plus font subsetting, do nearly all the meaningful work, since there's little to no image data to downsample. These documents can typically be compressed quite aggressively with no visible tradeoff at all, since the methods involved are almost entirely lossless.
Scanned books, forms, and archival documents. Downsampling is the dominant method by far, since every page is image data. Re-encoding strength should be chosen carefully based on how the document will be used — lighter for anything that needs to stay legible at a close zoom, stronger for anything that just needs to remain readable at normal viewing size.
Merged or heavily edited PDFs. Object deduplication and metadata stripping tend to deliver outsized results here, since these files accumulate redundant internal data over multiple rounds of editing that downsampling or re-encoding alone wouldn't touch.
Recognizing which category your document falls into is most of the battle — once you know that, you know roughly what kind of size reduction to expect and which settings actually matter for your specific file.
Frequently Asked Questions
What's the difference between lossy and lossless PDF compression?
Lossless compression shrinks data with nothing discarded, so it can be perfectly reconstructed — mainly used for text and internal file structure. Lossy compression discards some information to achieve smaller sizes, mainly used for images, and involves a real, if often minor, quality tradeoff.
Does downsampling always reduce image quality noticeably?
Not if it's done within a sensible range for how the image will be viewed. Downsampling to a resolution still well above what a screen displays is generally invisible; pushing it too low starts to visibly soften the image.
What is font subsetting and does it change how my document looks?
Font subsetting embeds only the specific characters actually used in a document instead of an entire font file. It's fully lossless and has no visible effect on the document's appearance.
Why does compressing a text document barely change the file size?
Because text-heavy files rely mostly on lossless methods like stream compression and font subsetting, which don't have as much raw data to remove compared to image-heavy files that benefit from lossy image compression.
Which compression method matters most for a scanned document?
Downsampling, since scanned pages are stored entirely as image data. The resolution setting chosen has the single biggest impact on both file size and visual quality for scans.
Can object and stream compression alone make a noticeable difference?
Yes, particularly for PDFs that have been edited or merged multiple times and accumulated redundant internal data, even if the visible content stays completely unchanged.
Is there one PDF compression method that works best for every document?
No. The best results come from combining methods appropriately based on content — heavier reliance on downsampling and re-encoding for image-heavy files, and heavier reliance on lossless structural methods for text-heavy ones.
Final Thoughts
PDF compression isn't one technique — it's a toolkit of several distinct methods, each suited to a different part of a document, and the best results come from applying the right combination rather than one blanket setting. Text and structure compress losslessly with almost no tradeoff. Images require a real, carefully calibrated lossy tradeoff. Knowing which methods matter for your specific document turns compression from a black box into something you can reason about.
If you'd rather skip the manual calculation, NanPDF's Compress PDF tool applies these methods automatically based on what's actually in your file, balancing downsampling, re-encoding, font subsetting, and structural cleanup so you get a meaningfully smaller PDF without having to pick each setting yourself.
Ready to try it yourself?
Compress PDF in seconds with nanPDF — free, secure, and no software to install.
Compress PDF