OCR Accuracy: What Actually Affects It and How to Get Better Results
Run the same OCR engine on two different documents and you can get wildly different results — one comes back nearly perfect, the other is riddled with garbled words and misplaced characters. It's tempting to conclude the software is inconsistent, but in most cases the tool didn't change at all. What changed was the input. OCR accuracy is less about which tool you use and more about a handful of specific, identifiable factors in the source document. Understanding them turns OCR from something that feels unpredictable into something you can reliably control.
What "OCR Accuracy" Actually Means
Before troubleshooting accuracy, it helps to be precise about what's being measured. OCR accuracy typically refers to how many characters, words, or lines the software correctly identifies compared to the original text. A tool with 99% character accuracy still means roughly one error every hundred characters, which sounds small until you realize a single-page document might have two or three thousand characters — meaning dozens of potential errors are still mathematically possible even at a high accuracy rate. This is why "accuracy" isn't a single fixed number for a given OCR engine; it shifts dramatically based on what you feed it.
The Biggest Factor: Scan Resolution
Resolution, measured in DPI (dots per inch), determines how much visual detail the OCR engine has to work with when distinguishing one character from another. At low resolution, letters that are naturally similar — a lowercase "l" and the number "1," or a "c" and an "e" — start to blur together into ambiguous shapes.
As a practical benchmark, 300 DPI is the widely recommended minimum for reliable OCR on standard printed text. Below roughly 150 DPI, accuracy tends to drop off noticeably, especially on smaller font sizes. Higher resolution generally helps, but there's a point of diminishing returns — going from 300 to 600 DPI rarely improves results much further, while it does increase file size and processing time. If you're scanning documents specifically for OCR, 300 DPI is usually the sweet spot.
Image Contrast and Cleanliness
OCR engines work by distinguishing text from background, and that distinction gets harder as contrast drops. A few common contrast problems come up repeatedly:
Faded or light print, common in older documents, carbon copies, or thermal receipts, gives the engine very little visual difference between the text and the page behind it.
Background noise, like scanned pages with visible paper texture, coffee stains, watermarks, or shadows from a poor scanning setup, can be misread as stray characters or can obscure real ones.
Colored or patterned backgrounds, such as forms with shaded boxes or letterhead graphics behind text, reduce the clean separation OCR relies on.
Show-through from the reverse side of thin paper, where text from the back of a page bleeds through visually on a scan, can confuse the engine into reading overlapping characters.
Where possible, scanning in black-and-white or grayscale mode rather than full color, and ensuring even lighting with no shadows, meaningfully improves contrast and therefore accuracy.
Skew and Page Orientation
A page that's scanned or photographed at even a slight angle throws off OCR more than most people expect. Text recognition generally assumes lines run roughly horizontal, and when a page is skewed, each line of text tilts progressively further from that expectation as it moves across the page, distorting how characters are grouped and read. Many modern OCR tools include automatic deskewing that straightens the page before recognition runs, which helps considerably, but severe skew — a page scanned at a sharp angle rather than a slight tilt — can still overwhelm that correction. Keeping the source page as straight as possible during scanning avoids relying on the correction step at all.
Font Type and Text Formatting
Not all fonts are created equal from an OCR engine's perspective.
Standard serif and sans-serif fonts used in most books, business documents, and reports are what OCR engines are most heavily trained on, and they tend to produce the highest accuracy.
Decorative, script, or highly stylized fonts deviate from standard letterforms enough that recognition accuracy drops, sometimes significantly, since the engine is essentially working outside its trained comfort zone.
Very small font sizes shrink the visual detail available per character, which compounds any resolution or contrast issues already present.
Dense formatting, like tightly kerned text, multi-column layouts, or text wrapped around images and tables, can cause the engine to misread word boundaries or process content in the wrong reading order, even if each individual character is recognized correctly.
Italics and bold text are generally handled well by modern engines, but heavily stylized combinations, like bold italic in a small size, add compounding difficulty.
Handwriting vs. Printed Text
This is one of the starkest accuracy gaps in OCR. Printed text follows standardized, repeatable letterforms that vary only in font choice — an "a" printed in one document looks essentially like an "a" printed in another. Handwriting has none of that consistency. Every person's handwriting is a unique combination of shape, spacing, slant, and stroke pressure, and even the same person's handwriting varies from one page to the next depending on speed and pen.
Neat, consistent print handwriting — think block capital letters written carefully — can sometimes be recognized reasonably well by modern OCR engines, particularly ones designed specifically for handwriting recognition. Cursive handwriting remains substantially harder, since letters connect and vary continuously rather than presenting as discrete, separable shapes. If a document is primarily handwritten, especially in cursive, it's realistic to expect meaningfully lower accuracy than you'd get from the same content printed, and manual review afterward becomes more important, not less.
Language and Character Set
OCR engines are trained on specific character sets, and selecting the correct language before processing has a real, measurable impact on accuracy. This matters for two reasons. First, different languages use different accented characters, ligatures, and symbols that the engine needs to know to look for. Second, many OCR systems use language models alongside pure character recognition — meaning they weigh how likely a given word is based on common word patterns in that language. If you run English-language recognition on a French document, the engine isn't just missing accent marks; it's also making worse guesses on ambiguous characters because it's checking them against the wrong dictionary of likely words entirely.
Realistic Accuracy Expectations by Document Type
It helps to calibrate expectations rather than assume every document will land at the same accuracy level.
Clean, digitally-printed documents (like a PDF exported from Word and then scanned, or a crisp modern printed page) typically see the highest accuracy, often in the high 90s percent for character recognition under good scanning conditions.
Older typewritten or photocopied documents tend to run somewhat lower, since irregular typewriter spacing, faded ink, and generational copy degradation all chip away at clarity.
Forms with mixed print and handwritten fields are inherently uneven — the printed labels will likely be recognized well, while handwritten answers in the blanks will vary widely in accuracy.
Low-quality photographed documents, especially those with shadows, glare, or an angle, will generally underperform a proper flatbed scan of the same page, sometimes substantially.
Dense technical documents with tables, formulas, or specialized symbols often see lower accuracy on the structured elements even when the surrounding prose text is recognized well, since layout complexity adds a separate challenge beyond character recognition itself.
How to Improve Accuracy Before Running OCR
- Step 1: Scan at 300 DPI or higher. This is the single highest-leverage change you can make, and it costs nothing beyond choosing the right scanner or camera setting.
- Step 2: Scan in grayscale or black-and-white rather than color for standard text documents, since it improves contrast between text and background without adding unnecessary visual noise for the engine to process.
- Step 3: Keep the page flat and the shot straight-on. Avoid curled corners, folds, and camera angles that introduce skew or perspective distortion.
- Step 4: Ensure even, shadow-free lighting. Uneven lighting creates artificial contrast variation across the page that can confuse recognition in the darker areas.
- Step 5: Select the correct language setting before processing. Don't leave this on a default if the document isn't in that language — it meaningfully changes recognition quality.
- Step 6: Crop out irrelevant background if you're photographing a document rather than using a flatbed scanner, so the engine focuses only on the actual content.
How to Improve Results After Running OCR
Even with a clean scan, it's worth building a habit of reviewing OCR output rather than assuming perfection, especially for anything with legal, financial, or medical stakes.
Spot-check numbers and names first. These are the highest-consequence errors, since a misread digit in an invoice total or a misspelled name in a legal document carries more risk than a minor typo in a sentence of prose.
Re-scan and re-run rather than manually fixing a bad source. If accuracy came back poor because of a genuinely bad scan, it's often faster to rescan properly and reprocess than to manually correct dozens of errors from a flawed source image.
Use a tool that preserves the original image alongside the text layer, like NanPDF's OCR PDF tool, so you can always cross-reference the recognized text against the actual scanned page if something looks questionable.
Frequently Asked Questions
What DPI should I scan at for the best OCR accuracy?
300 DPI is the generally recommended standard for printed text. Going higher rarely improves results much further, while going much lower, especially under 150 DPI, tends to noticeably reduce accuracy.
Why is my OCR accuracy worse on an old document than a new one?
Older documents often combine several accuracy-reducing factors at once: faded ink, inconsistent typewriter spacing, physical degradation from age, and lower-quality reproduction if they've been photocopied multiple times over the years.
Does color scanning improve or hurt OCR accuracy?
For standard black-text-on-white-page documents, grayscale or black-and-white scanning typically improves contrast and therefore accuracy compared to full color, which can introduce unnecessary visual noise.
Can OCR accurately read handwriting?
It depends heavily on the handwriting style. Neat, consistent print handwriting can be recognized reasonably well by modern engines, but cursive and inconsistent handwriting remain considerably harder and should be manually reviewed.
Does selecting the wrong language really make a noticeable difference?
Yes. OCR engines often use language-specific models to predict likely words alongside raw character recognition, so processing a document in the wrong language setting can reduce accuracy well beyond just missing accented characters.
Is there a way to know how accurate my OCR result is without checking every word?
Spot-checking a representative sample — particularly numbers, names, and any dense or unusual formatting — is a practical middle ground between blind trust and reviewing every single character.
Why does the same OCR tool give different accuracy on different documents?
Because accuracy depends far more on the input than the tool itself. Resolution, contrast, skew, font, handwriting, and language all vary from document to document, and each one independently affects how well the engine can recognize the text.
Final Thoughts
OCR accuracy isn't a fixed property of the software you choose — it's largely a function of what you feed into it. A clean, well-lit, high-resolution scan of standard printed text in the correct language will consistently outperform a blurry, skewed photo of handwritten notes, regardless of which engine processes them. Once you know which levers actually move the needle, getting reliable results stops being guesswork. If you want to see how your own documents perform, NanPDF's OCR PDF tool processes files directly in the browser, so you can test a scan, review the results, and adjust your approach in minutes rather than committing to a workflow blind.
Ready to try it yourself?
Try OCR PDF in seconds with nanPDF — free, secure, and no software to install.
Try OCR PDF