---
title: "Where are the regions?"
url: https://sharadbapat.com/experiments/where-are-the-regions/
description: "Everything a PDF page draws, each with its box, read straight from the file. On 967 test pages the boxes hold every pixel of ink."
published: 2026-10-03
author: Sharad Bapat
---

# Where are the regions?

A PDF page is a list of drawing instructions: text, images, shapes and annotations. I built a small reader that lists everything a page draws, with a box for each, straight from the file. Then I checked that its boxes hold all the ink when the page is drawn.

[Try the demo](https://sharadbapat.com/experiments/where-are-the-regions/demo/): drop in a PDF and see every region boxed on its page.

[Code on GitHub](https://github.com/sharad-bapat/where-are-the-regions)

- **Made:** Oct 2026
- **Built with:** Rust, compiled to WebAssembly; pdf.js draws the page in the demo
- **Status:** Working; there's a demo

## The question

This is the fourth tool in a series. [What is this file?](https://sharadbapat.com/experiments/what-is-this-file/) says what an upload really is, [Scan or text?](https://sharadbapat.com/experiments/scan-or-text/) says whether a PDF needs OCR, and [Where is the text?](https://sharadbapat.com/experiments/where-is-the-text/) finds every word with its box. Words aren't all a page holds, though. There are pictures, scanned pages, tables drawn as lines, charts, logos, text drawn in white or hidden behind a clip. Before deciding which parts of a page to send to OCR, I wanted a complete map of what's on it.

So the question was whether a reader with no model and no rendering could list every mark a page draws, each with a box, and be checkably complete: no ink on the page outside its boxes.

## How it works

It follows each page's drawing instructions in order and records every mark: each line of text, each image, each drawn shape (grouped into clusters of touching paths) and each annotation. Each one gets a box on the page as it's displayed, its place in the drawing order, and flags for the ways a reader might not see it: cut away by a clip, off the edge of the page, painted in white, or invisible text like an OCR layer over a scan.

Text boxes come from the font. A word's box runs along its glyphs' advance widths, from the font's descent to its ascent, and where a glyph's outline reaches further (an italic f, a capital taller than the font says) the word gets a second, larger ink box. The outlines come from the font program embedded in the file, or for fonts that aren't embedded, from the metrics of the standard font a reader would draw instead.

On top of that map sits a second layer that guesses what things are. An image is judged from its own pixels: a photo, a scanned page of text, a graphic or blank. A cluster of drawn shapes is judged from its lines and fills: a rule, a border, a table grid, a fill, a chart or diagram, or text drawn as letter outlines. Each guess comes with a confidence and the reasons behind it. The map should be exact. The kinds are guesses, so they're scored by how often their confident answers are right.

## How I tested it

The main check doesn't trust the reader at all. The ink test draws each page with PyMuPDF 1.24.9 at 150 dpi and counts the dark pixels that fall outside every box, with a point of margin. It also reports boxes that hold no ink, which point at marks the map has wrong. Beside it are constructed sets with answers known by construction: 200 pages whose every mark has a known box, 600 pages of images, and 240 pages of drawings of each kind.

Real pages came from [govdocs1](https://digitalcorpora.org/corpora/file-corpora/files/): 278 files from thread 003 for tuning (967 labelled pages), and 235 from thread 004 held out (704 pages). The image labels came from Tesseract run on each image's pixels, and Claude checked a sample of them by eye. For the drawings, Claude labelled 150 clusters per thread by eye from contact sheets, without the predictions in view.

I tuned the rules on the tuning data only, froze them by a hash of the source, and ran them once on the held-out data against targets set beforehand. Problems the held-out run found were fixed afterwards in their own commits, and I report those results apart from the held-out ones.

## Results

- **1,015 of 1,015**: held-out constructed marks boxed within 1 pt
- **967 of 967**: tuning pages with every pixel of ink inside a box

On the held-out pages, with the rules frozen, 703 of the 704 had at least 99.5% of their ink inside a box, against a target of every page. The one that missed (99.26%) drew its text in Type 3 fonts whose widths say zero, and the next lowest (99.81%) used a Symbol font the file named but didn't embed. With those two fixed, and then glyph outlines added for every kind of font program, 703 of the 704 held-out pages are at 100%, and the last one misses 3 pixels.

To see whether that means anything, I ran the same ink test on maps built from other readers' boxes, on the 967 tuning pages:

| Reader | Pages at 100% | Worst page | Ink outside every box |
| --- | --- | --- | --- |
| this reader | 967 | 100% | 0 pixels |
| PyMuPDF get_bboxlog | 966 | 56.42% | 710,374 pixels |
| pdfplumber 0.11.9 | 722 | 36.55% | 781,588 pixels |
| PyMuPDF's usual calls (get_text, get_image_info, get_drawings) | 720 | 14.88% | 2,076,528 pixels |

The test favours bboxlog, which is MuPDF reporting on MuPDF's own drawing of the page. It loses one page, where the background is filled with a repeating pattern and bboxlog boxes it by a single tile. The calls most people use miss whole marks on 37 files, and pdfplumber on 49. On this test the reader is level with MuPDF's own view of the page, one page ahead, from a parser that doesn't use MuPDF, and it adds the flags and the drawing order.

## The kinds

The image kinds met their held-out targets. On the constructed set, 147 of 150 text images were found, and none of the 88 photos, logos, blanks and rules was called text. Across real and constructed images, 91.4% of the kinds given with a confidence of 0.9 or more were right.

The drawing kinds didn't. The constructed set was right 117 of 117, but on the real held-out pages only 66.7% of the kinds given at 0.9 or more were right, against a target of 90%. On the tuning pages the same figure was 90.2%, so the rules fitted those 46 files and didn't carry over. Rules and table grids held up best; charts were the weak spot, and 10 of the 12 charts called table grids were framed plots. I'd treat the drawing kinds as experimental, and the demo marks them that way.

## Speed and size

|  | 003, 7,633 pages | 004, 4,211 pages |
| --- | --- | --- |
| native | 2.03 ms | 1.69 ms |
| WebAssembly in Node | 1.90 ms | 2.55 ms |

That's the median time a page, timed from the file's bytes in memory, in one session on my laptop. A few files set the totals: one 76-page file in 004 takes about 1.3 seconds a page, and I haven't profiled it. The WebAssembly build is 772 KB, 315 KB compressed, and its output matches the native build's byte for byte on all 1,553 test files.

## Limits

Only PDFs are read; image files as input aren't handled yet. There's no reading order and no grouping beyond word, line and touching shapes, and overlap isn't tracked, so a mark painted over by a later one still counts as visible. For fonts the file doesn't embed, glyph boxes use Adobe's metrics, while MuPDF draws those fonts with URW's Nimbus fonts, which differ by a few units on some glyphs. JPX and JBIG2 images aren't decoded, so they get no kind. The ink test's judge is MuPDF, so where MuPDF draws a page differently from other renderers, the test inherits that.

## What didn't work

The ink test itself had two faults at first. PyMuPDF reports boxes on the unrotated page but draws it rotated, so on turned pages it blamed the wrong marks for missed ink, and it counted the space between words as text. Both had to be fixed before its numbers meant anything.

The first boxes stopped at the advance width and the font's stated ascent, which is what most text tools report. That left italic overhang and tall capitals outside on 180 pages. Reading glyph outlines fixed it in five steps: embedded TrueType and CFF fonts, fonts that aren't embedded, Type 1 fonts, CFF glyphs ttf-parser won't read, and glyphs that cross the page edge. Two of my own bugs turned up along the way. A Type 1 font's subroutines weren't read at all, because the parser stopped at the word "array". And every bold italic Arial got the boxes of plain bold, because two fonts with identical width tables were told apart by where those tables sat in memory.

Some files carry the descent as a positive number, which the spec says should be negative, and my first parser clamped it to zero, so every word box stopped at the baseline. I also tuned the drawing kinds on 46 files, which wasn't enough. They didn't hold on new pages.

## Findings about other tools

The checks turned up some behaviour in other software. None of it is reported upstream yet.

- PyMuPDF boxes a shape filled with a tiling pattern by one tile at the pattern's origin, not by the area it paints.
- PyMuPDF's word extraction leaves out glyphs whose codes decode to control characters, so pages in some fonts lose most of their words.
- pdfminer.six, which pdfplumber uses, ignores the shading operator, so shaded areas aren't listed.
- ttf-parser gives up on any CFF glyph that uses the deprecated dotsection operator, which Adobe's spec says to treat as a no-op. About 2% of the glyph programs in govdocs1's embedded CFF fonts use it.
- MuPDF doesn't draw a round-capped point given as a single closed subpath, and pdfium draws zero-length lines with butt or square caps, which the spec says draw nothing.

## Credits

- govdocs1 from Digital Corpora (Garfinkel et al., 2009), public domain, for the real pages. None of them is served here. The four samples in the demo are made-up files generated for this page.
- [pdf.js](https://mozilla.github.io/pdf.js/) by Mozilla (Apache-2.0) draws the page in the demo, and its licence is bundled with it.
- PyMuPDF for the ink test's drawing, PyMuPDF and pdfplumber for the comparison, Tesseract for the image labels, and fontTools for checking glyph boxes. ttf-parser reads embedded TrueType and CFF fonts, zune-jpeg decodes JPEG images, and the standard-font metrics are Adobe's core-14 AFM files.
- Claude labelled the drawing clusters and checked the image labels, as described above.
