Skip to content
AITroveRead. Build. Understand.
Make this comfortable

OCR text identity: geometry, reading order and offsets

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Extracted text is a hypothesis tied to an image region. Preserve that tie before using words in search, entity extraction or decisions.

Keep the page as evidence

For each line or token, store document revision, page number, bounding region, OCR text and recognizer version. Reading order is not guaranteed by y-coordinate alone: columns, tables and footnotes can interleave. Preserve the image and a structured block graph; a flattened paragraph is a derived view. If an agent corrects the text, keep both OCR output and the reviewed revision. Span offsets must refer to a named representation.

Do not equate characters across views

OCR may turn “O” into “0,” omit punctuation or split a word at a line break. Downstream token offsets in corrected text cannot be applied to the OCR string without an alignment map. Record the original block IDs that produced each normalized span. For tables, keep row and column relationships rather than joining every cell into one sentence. A total beside a date should not become an invented amount-date phrase.

Use geometry to review errors

Show a reviewer the source crop alongside the extracted field. A text-only correction can hide whether the wrong number came from a neighboring line. Validate page and box bounds before cropping. A rotated scan, faint print or mixed handwriting should be flagged for manual inspection when field confidence is low. The field review policy turns these signals into an abstain decision.

Evaluate the downstream task

Character error rate is useful, but it does not tell whether a specific invoice amount or ID is wrong. Measure exact-field accuracy, misplaced spans, table association errors and correction time by document type. Keep scans from one document family in one split. The intake project carries page evidence through to a reviewed field.

Implementation

python
def make_ocr_span(page_text, start, end, page_number, box):
    if not 0 <= start < end <= len(page_text):
        raise ValueError("span falls outside OCR text")
    if page_number < 1 or len(box) != 4 or any(not 0 <= value <= 1 for value in box):
        raise ValueError("invalid page or normalized box")
    return {"text": page_text[start:end], "start": start, "end": end,
            "page": page_number, "box": tuple(box)}

span = make_ocr_span("Invoice ZX-47 total 2847", 8, 13, 2,
                     (0.14, 0.21, 0.31, 0.25))
assert span["text"] == "ZX-47" and span["page"] == 2

Performance and operating cost

The validation and slice copy cost O(length of span) time and space. Layout analysis and image recognition dominate compute, while page crops increase evidence storage. Do not discard geometry merely to save a small metadata field; without it, reviewers cannot trace a wrong value to its source region.

Common Mistakes

  • Treating flattened OCR text as the only document representation.
  • Applying corrected-text offsets to raw OCR output.
  • Joining table cells in arbitrary reading order.
  • Approving a critical number without viewing its source crop.

Read next

Continue the workflow: Table evidence: cell coordinates, headers and reading order.

ai-data
natural-language-processing
Storage details