What needs OCR?

Most PDFs made on a computer carry their text, but parts of a page can exist only as pixels. I built a router that finds those parts from the page map, sends only them to Tesseract, and merges what it reads with the file's own words. On held-out government PDFs it OCR'd 5.5% of the page area instead of 98.8%, and its word error was lower than OCR'ing every page.

Try the demo: drop in a PDF and see which parts of each page would go to OCR, and why.

Code on GitHub

Made
Built with
Rust on the where-are-the-regions page map, compiled to WebAssembly; Tesseract 5.4.0 for the OCR; pdf.js draws the page in the demo
Status
Working; there's a demo

The question

This is the fifth tool in a series. Scan or text? says whether a PDF needs OCR at all, Where is the text? finds every word with its box, and Where are the regions? lists everything a page draws. Scan or text? answers once per file, and a file that has text can still hide some that it doesn't: a signed page pasted in as a scan, a chart exported as a picture, a logo, a title converted to outlines, or a font whose text copies out as junk.

The usual fix is to OCR every page. That catches those parts, but it costs a lot, and it reads again the text the file already has exactly, and misreads some of it. So the question was whether the page map could say which parts of a page need OCR, and whether reading only those gives text as good as OCR'ing every page for much less.

How it works

The router reads the page map and makes three kinds of route. Every image the page draws gets a score: the map's score for its size and place on the page, times the evidence from its own pixels that it holds text. At 0.1 or more it goes to OCR, and an image already covered by invisible text (an earlier OCR) is left alone. Drawings go when the map calls them letters drawn as outlines, grouped into blocks like lines of text, or when a chart or table drawing has at least three letter shapes in it. And when at least half of a font's visible words don't decode (private-use or control characters, U+FFFD, or codes with no Unicode at all), those words are grouped into blocks and sent too.

Each route is cropped from the page with 6 pt of margin, rendered by PyMuPDF, turned so its text runs left to right, and read by Tesseract. Images are read at their own resolution, kept between 300 and 400 dpi, and everything else at 300. The merge keeps every word the file shows, drops an OCR word that lands on one of them, and swaps the undecodable words for what OCR read in their place. OCR words under a confidence of 0.5 are dropped as well (that rule came later, below). Each word comes out with its box, where it came from and a confidence.

Routing doesn't render anything or run OCR, so the same file always gives the same routes, and it runs in the browser in the demo. OCR doesn't run there. The demo's three samples carry what Tesseract read in their crops, made beforehand, and the merge runs live on those.

How I tested it

The constructed set is 600 one-page PDFs rebuilt from the NDAs in ContractNLI, where the right answer is known by construction. Each page gets a rasterised crop of another contract pasted in, or is turned into a full scan, or has a crop's words redrawn as outlines or with a garbage ToUnicode map. The negatives are a crop that already has an OCR layer, a photo, a logo, a blank sheet, or the page left as it was. Tesseract found 99.1% of the truth words when checking it.

Real PDFs have no truth. I used pages from govdocs1 that have a usable text layer: thread 005 for tuning (668 pages), 006 held out (636 pages) and 007 kept for testing a change made after the held-out run (613 pages). A reference was built for each page without the router: the file's visible words that decode, plus the words Tesseract reads at 400 dpi with a confidence of 85 or more that don't sit on a file word. Claude checked 20 pages of it by eye, and 274 of the 299 reference words looked at were right (92%). It also misses many small labels on maps and charts. So the numbers on real pages are agreement with the reference, not accuracy.

Words are matched one to one by their letters and digits and by position, and word error is missing plus extra words over the truth's word count. There are three baselines: OCR every page at 300 dpi, OCR only the pages with no text layer (as ocrmypdf's --skip-text does), and the file's text alone. I tuned on the tuning data only, froze the source by hash with where-are-the-regions pinned at its commit, and ran the held-out data once against targets set beforehand.

Results

210 of 210held-out constructed regions that need OCR were routed
94% lesspage area sent to OCR on held-out government pages

On the constructed held-out split, the 210 routed regions held 99.9% of their words, and none of the 30 already-OCR'd crops or 45 photos, logos and blank sheets was sent to OCR. The whole pipeline, held out, with the source frozen:

The whole pipeline on the held-out data, against OCR'ing every page and the file's text alone
SetMethodRecallWord errorPage area OCR'dTesseract seconds
constructed, 300 pagesrouted99.70%0.49%21.95%299
every page98.93%1.76%100%922
file text76.49%23.52%0%0
govdocs1 006, 636 pagesrouted99.19%2.09%5.50%118
every page96.90%6.17%98.78%1,946
file text98.53%1.68%0%0

All three targets were met. Routing found at least 98% of the regions that need OCR (all of them), the pipeline's word error was no more than a point worse than OCR'ing every page (it was 1.27 points better), and on real pages it sent at least 70% less area to OCR with no more than a point more word error (94% less area, 4.08 points less word error). OCR'ing every page does worse on word error because it reads the file's own words again, and Tesseract gets some of them wrong.

The table also shows where it fell short. On 006 the pipeline's word error was higher than the file's text alone, 2.09% against 1.68%, and most of the extra words were OCR words that Tesseract wasn't sure of: 2,087 of 2,699 had a confidence under 85. After the held-out run I added a cutoff to the merge, chosen on the tuning data by a rule written down before testing (the lowest summed word error over the constructed tuning split and 005, keeping 98% of the regions). It came out at 0.5. 006 had been used, so the test was one run on 007:

The confidence cutoff tested on govdocs1 thread 007
govdocs1 007, 613 pagesRecallWord errorPixel-only words foundPage area OCR'dTesseract seconds
routed, cutoff 0.598.97%1.42%83.60%2.35%92
routed, no cutoff98.98%1.65%83.87%2.35%92
every page96.12%6.90%91.91%98.25%1,738
file text98.32%1.85%9.56%0%0

With the cutoff the pipeline's word error is lower than the file's text alone, for 0.01 points of recall. Pixel-only words are the reference's words that exist only as pixels, the ones this tool is for. The router finds about 78 to 84% of them, where OCR'ing every page finds 88 to 92%, for about a twentieth of the Tesseract time. That gap is the main thing left to work on.

Limits

The pixel-only words it misses fall into a few groups. Letters drawn as outlines up the side of a chart aren't found, because the page map only looks for letters in horizontal runs. Logos whose pixels score at the map's floor for text are skipped, along with the photos and blank sheets that score the same. A crop of a photo or a map reads less than the whole page at 400 dpi does. Bullets stored as private-use characters, in a font of their own, are sent to OCR as undecodable text. Tesseract is the only engine I tried, in English only, and I haven't timed the routing itself. The WebAssembly build is 861 KB, 350 KB compressed, and its output matches the native build's byte for byte on 852 files.

What didn't work

The reference for the real pages took four rebuilds, each after looking at the pages. The first removed the text layer and read what was left, but removal left text behind, and on one file the whole body survived as "pixel" words. At a confidence of 60, chart and map marks came through as words like "tLe" and "=a", so the cutoff went to 85. A text layer that decodes to Latin-1 letters mixed with C1 control characters passed as good text, and the first rule written to catch layers like it flagged 5 good layers out of 7. And pages with /Rotate were read sideways. The crop code had the same rotation bug, so the fix went into the pipeline too.

Letters drawn as outlines first came out as dozens of tiny routes, one for each cluster of letters, until they were grouped into blocks. Inline images had no pixels to judge, so where-are-the-regions had to learn to decode them. The first held-out run on 006 crashed before writing anything: when the same crop came up twice, two threads shared a cache entry, and one read the other's half-written file. The fix changed no OCR word, and the run was made again after relocking.

On the constructed held-out split, one page got a real wrong route: two bullets stored as a private-use character, the only two words in their font, were sent to OCR as an undecodable layer. And the held-out run on 006 showed the pipeline losing to the file's text on word error, which is what the cutoff fixed, measured on pages nothing was chosen on.

Credits