What needs OCR? Demo
pdf.js draws the page. Every part of it whose text only exists as pixels is boxed on top and listed beside it, with the reason the router gives. The three samples also show the merged words: the file's own text plus what Tesseract read in their crops.
This demo needs JavaScript. It reads a PDF you choose in this page, with nothing uploaded, draws each page with pdf.js, and boxes every part the router would send to OCR.
Orange boxes go to OCR. Grey dashed boxes are images it skips, and green ones are already covered by an earlier OCR. In the words view, OCR words are orange and the file's own words are faint blue. The write-up explains how it works and how it was tested.