Skip to content
AITroveRead. Build. Understand.
Make this comfortable

OCR field review: confidence, consistency and abstention

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A high line score is not a guarantee that a critical amount is right. Review policy belongs to each field and workflow.

Calibrate by field and source

OCR systems may return confidence for words, lines or extracted fields. These are different events. A line can be readable while a table relationship is wrong; a field can have a plausible shape while the digit is mistaken. Collect reviewed examples by scan quality and document layout, then estimate field correctness at each score range. Use a separate threshold for critical values such as invoice amount or bank reference. Do not describe a raw score as a calibrated probability without evidence.

Combine independent checks

Validate syntax, expected value ranges, duplicate occurrences and cross-field relationships. A checksum can reject an impossible identifier, but passing one does not prove the value was read from the right box. When two occurrences disagree, show both source crops. Never overwrite the extracted value with a guessed correction without preserving the original. Geometry and offset identity let a reviewer inspect the actual image.

Route uncertainty to people

Return accept, review or reject with reason codes. Reject an invalid page reference or missing evidence; review low-confidence critical fields and conflicting totals. For a noncritical searchable phrase, a lower threshold may be tolerable if it cannot trigger an external action. Track reviewer changes and false accepts separately. A conservative system can be useful even when many documents need review, if it reduces transcription time without automating wrong amounts.

Monitor incoming documents

Watch field accuracy, review fraction, time to resolution and error rate by template, scanner and language. New forms or rotated photos change the distribution. Shadow-test a new OCR model against the same reviewed pages before switching. A removed document should remove its extracted text, crops and downstream search records. The intake project bundles these controls.

Implementation

python
def decide_ocr_field(field, critical_thresholds):
    if not field.get("source_box") or not field.get("source_revision"):
        return "reject", "missing-source-evidence"
    if not field["syntax_valid"]:
        return "review", "invalid-format"
    threshold = critical_thresholds.get(field["kind"], 0.72)
    if field["confidence"] < threshold or field.get("conflict", False):
        return "review", "uncertain-field"
    return "accept", "reviewed-policy"

amount = {"kind": "amount", "confidence": 0.91,
          "syntax_valid": True, "source_box": (0.2, 0.3, 0.4, 0.4),
          "source_revision": "scan-47"}
assert decide_ocr_field(amount, {"amount": 0.95})[0] == "review"

Performance and operating cost

The gate is O(1) time and space per field. Maintaining calibrated thresholds takes reviewed documents from each important layout and source slice. Human review is the main operating cost; report false accept rate against review volume, rather than claiming that a single higher threshold is always better.

Common Mistakes

  • Applying a line-level score as a field-level guarantee.
  • Accepting a high-score amount from the wrong table row.
  • Discarding original OCR output after an agent correction.
  • Setting one threshold for every field and document template.

Read next

Continue the workflow: Extraction validation: repair limits and conflicting values.

ai-data
natural-language-processing
Storage details