A high line score is not a guarantee that a critical amount is right. Review policy belongs to each field and workflow.
OCR field review: confidence, consistency and abstention
Calibrate by field and source
OCR systems may return confidence for words, lines or extracted fields. These are different events. A line can be readable while a table relationship is wrong; a field can have a plausible shape while the digit is mistaken. Collect reviewed examples by scan quality and document layout, then estimate field correctness at each score range. Use a separate threshold for critical values such as invoice amount or bank reference. Do not describe a raw score as a calibrated probability without evidence.
Combine independent checks
Validate syntax, expected value ranges, duplicate occurrences and cross-field relationships. A checksum can reject an impossible identifier, but passing one does not prove the value was read from the right box. When two occurrences disagree, show both source crops. Never overwrite the extracted value with a guessed correction without preserving the original. Geometry and offset identity let a reviewer inspect the actual image.
Route uncertainty to people
Return accept, review or reject with reason codes. Reject an invalid page reference or missing evidence; review low-confidence critical fields and conflicting totals. For a noncritical searchable phrase, a lower threshold may be tolerable if it cannot trigger an external action. Track reviewer changes and false accepts separately. A conservative system can be useful even when many documents need review, if it reduces transcription time without automating wrong amounts.
Monitor incoming documents
Watch field accuracy, review fraction, time to resolution and error rate by template, scanner and language. New forms or rotated photos change the distribution. Shadow-test a new OCR model against the same reviewed pages before switching. A removed document should remove its extracted text, crops and downstream search records. The intake project bundles these controls.
Implementation
def decide_ocr_field(field, critical_thresholds):
if not field.get("source_box") or not field.get("source_revision"):
return "reject", "missing-source-evidence"
if not field["syntax_valid"]:
return "review", "invalid-format"
threshold = critical_thresholds.get(field["kind"], 0.72)
if field["confidence"] < threshold or field.get("conflict", False):
return "review", "uncertain-field"
return "accept", "reviewed-policy"
amount = {"kind": "amount", "confidence": 0.91,
"syntax_valid": True, "source_box": (0.2, 0.3, 0.4, 0.4),
"source_revision": "scan-47"}
assert decide_ocr_field(amount, {"amount": 0.95})[0] == "review"
Performance and operating cost
The gate is O(1) time and space per field. Maintaining calibrated thresholds takes reviewed documents from each important layout and source slice. Human review is the main operating cost; report false accept rate against review volume, rather than claiming that a single higher threshold is always better.
Common Mistakes
- Applying a line-level score as a field-level guarantee.
- Accepting a high-score amount from the wrong table row.
- Discarding original OCR output after an agent correction.
- Setting one threshold for every field and document template.
Read next
- OCR text identity: geometry, reading order and offsets
- Project: review OCR fields in a document-intake queue
- Decode entities and evaluate exact spans, not token accuracy
- PII detection and redaction on original text offsets
- Text classification evaluation: inspect slices and allow abstention
Continue the workflow: Extraction validation: repair limits and conflicting values.
