Extracted text is a hypothesis tied to an image region. Preserve that tie before using words in search, entity extraction or decisions.
OCR text identity: geometry, reading order and offsets
Keep the page as evidence
For each line or token, store document revision, page number, bounding region, OCR text and recognizer version. Reading order is not guaranteed by y-coordinate alone: columns, tables and footnotes can interleave. Preserve the image and a structured block graph; a flattened paragraph is a derived view. If an agent corrects the text, keep both OCR output and the reviewed revision. Span offsets must refer to a named representation.
Do not equate characters across views
OCR may turn “O” into “0,” omit punctuation or split a word at a line break. Downstream token offsets in corrected text cannot be applied to the OCR string without an alignment map. Record the original block IDs that produced each normalized span. For tables, keep row and column relationships rather than joining every cell into one sentence. A total beside a date should not become an invented amount-date phrase.
Use geometry to review errors
Show a reviewer the source crop alongside the extracted field. A text-only correction can hide whether the wrong number came from a neighboring line. Validate page and box bounds before cropping. A rotated scan, faint print or mixed handwriting should be flagged for manual inspection when field confidence is low. The field review policy turns these signals into an abstain decision.
Evaluate the downstream task
Character error rate is useful, but it does not tell whether a specific invoice amount or ID is wrong. Measure exact-field accuracy, misplaced spans, table association errors and correction time by document type. Keep scans from one document family in one split. The intake project carries page evidence through to a reviewed field.
Implementation
def make_ocr_span(page_text, start, end, page_number, box):
if not 0 <= start < end <= len(page_text):
raise ValueError("span falls outside OCR text")
if page_number < 1 or len(box) != 4 or any(not 0 <= value <= 1 for value in box):
raise ValueError("invalid page or normalized box")
return {"text": page_text[start:end], "start": start, "end": end,
"page": page_number, "box": tuple(box)}
span = make_ocr_span("Invoice ZX-47 total 2847", 8, 13, 2,
(0.14, 0.21, 0.31, 0.25))
assert span["text"] == "ZX-47" and span["page"] == 2
Performance and operating cost
The validation and slice copy cost O(length of span) time and space. Layout analysis and image recognition dominate compute, while page crops increase evidence storage. Do not discard geometry merely to save a small metadata field; without it, reviewers cannot trace a wrong value to its source region.
Common Mistakes
- Treating flattened OCR text as the only document representation.
- Applying corrected-text offsets to raw OCR output.
- Joining table cells in arbitrary reading order.
- Approving a critical number without viewing its source crop.
Read next
- OCR field review: confidence, consistency and abstention
- Project: review OCR fields in a document-intake queue
- Entity spans: align annotations to the original text
- Unicode and tokenization: preserve meaning at the text boundary
- Decode entities and evaluate exact spans, not token accuracy
Continue the workflow: Table evidence: cell coordinates, headers and reading order.
