Test results that stopped being retyped
Results arrived as images and were keyed in by hand. We automated the read — and kept every figure traceable to the page it came from, because accreditation requires it.
- Sector
- Food testing laboratory, Europe
- Service
- Document AI — Beyond OCR
- Deployment
- Private, no public LLM
- Constraint
- Accreditation requires provenance for every reported figure
The client is a food testing laboratory in Europe. Samples arrive from producers and retailers, are analysed, and results are reported back. The data is client-confidential, and the lab works to accreditation requirements — every figure it reports has to be traceable to its origin.
Results came off instruments and out of partner processes as images. Someone then read those images and typed the values into the lab system. It was slow, it consumed qualified staff time on transcription, and it introduced a category of error that is particularly costly here: a mistyped figure becomes a reported result.
The obvious fix — automated extraction — was also the obvious risk. An extraction pipeline that is usually right creates a review burden and an accreditation problem at the same time.
Automated reading of the result images into structured, validated records — designed so that the lab can defend every value it publishes.
Values and tables read from images whose layout varies by instrument and source, with the position on the page retained alongside each field.
Range checks and consistency rules catch what a model cannot know — a value a test would never produce, a figure that contradicts another on the same sheet.
Every extracted value carries a score. Low-confidence values never enter the lab system silently — they go to review.
The reviewer sees the extracted value next to the source image, so a check takes seconds rather than a reread of the whole sheet.
The source image is retained against every figure, with a record of what was extracted, what was flagged, and who confirmed it. This is the part that made the automation acceptable rather than merely useful.
On privately hosted models. No test data, sample reference or client identifier leaves the approved environment.
For a lab holding confidential results on behalf of producers and retailers, sending that content to a third-party model API was not an option worth discussing. Deciding this first shaped the architecture rather than constraining it late.
How we deploy privately →In regulated work, provenance is not a feature. It is the requirement.
Extraction alone would not have passed. An accredited lab has to show where every number came from, and a pipeline that produces values without a defensible trail moves the problem rather than solving it.
The retained source image, field-level confidence and review record are what turned automation into something the lab could put its name to.
Send us the document your team keys in by hand.
A discovery session is a working conversation, not a demo. Bring a handful of real examples, including the bad scans.