Build a document queue that keeps image evidence, extracted values and every correction tied to the same source revision.
Project: review OCR fields in a document-intake queue
Define what may be automated
An intake service reads an uploaded invoice and proposes vendor, invoice ID and amount. It may populate a review form, but it does not approve a payment. Store a source revision for the image, block geometry, raw OCR output and corrected field. Let a reviewer open the exact crop and resolve conflicts. The final accepted value must carry the reviewer decision and original evidence reference.
Assemble a representative audit
Sample different templates, scans, photos, rotated pages, low contrast and multiple languages. Have reviewers label field values and source boxes, including missing fields and wrong-row decoys. Group repeated vendor templates and near-identical documents across splits. Measure exact-field accuracy, abstention and review time; a character-level score alone misses financially important digit errors. The geometry contract preserves traceability.
Implement the handoff
Validate OCR block bounds, extract field candidates and run syntax and consistency checks. Propose an accept only if field-specific audit policy permits it; otherwise display original crop and candidate to a reviewer. The field gate separates low confidence from missing source evidence. Record a correction as a new reviewed value, never as a silent mutation of raw OCR.
Release and retire
Compare model versions on a frozen set before shadow rollout. Monitor errors by layout and upload channel. If a new template produces conflicting totals, route that slice to review. On source deletion, remove raw image, OCR text, crops, embeddings and derived index records under the retention policy. Keep only permitted aggregate metrics after deletion.
Implementation
def record_field_review(raw_value, proposed_value, source_revision,
reviewer_id, decision):
if decision not in {"accept", "correct", "reject"}:
raise ValueError("unknown review decision")
if not source_revision or not reviewer_id:
raise ValueError("review requires evidence and reviewer")
accepted = proposed_value if decision in {"accept", "correct"} else None
return {"raw": raw_value, "accepted": accepted,
"source_revision": source_revision, "reviewer": reviewer_id,
"decision": decision}
review = record_field_review("2847", "2847", "scan-47", "agent-82", "accept")
assert review["accepted"] == "2847" and review["raw"] == "2847"
Performance and operating cost
The review record is O(1) for field metadata plus O(value length) copying. Image processing, storage and human review dominate. Retaining crops alongside images adds bytes but reduces the cost of tracing a contested field. Measure correction time and false accepted amounts before optimizing extraction throughput.
Common Mistakes
- Using OCR text alone as payment approval.
- Overwriting raw evidence with a corrected value.
- Auditing only clean scans from one template.
- Deleting a source image while leaving extracted text and vectors searchable.
Read next
- OCR text identity: geometry, reading order and offsets
- OCR field review: confidence, consistency and abstention
- PII detection and redaction on original text offsets
- Audit redacted text flows, retention and re-identification risk
- Project: ship an auditable support-entity extractor
Continue the workflow: Project: preserve text identity in support intake.
Continue the workflow: Project: answer from support tables with cell evidence.
