Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Schema-bound extraction: typed fields and exact source spans

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Extract a document into typed fields whose values can be traced to exact source spans; a well-formed object alone is not evidence.

Define fields before decoding

A change request may mention a service, requested window and approval status in different paragraphs. Define each field’s meaning, type, cardinality and allowed null state before a model reads the text. A “window” can mean start time, end time or a duration, so one loose string field is inadequate. Keep schema version with every extraction. Time anchors and zones are needed before a local time can become a scheduled instant.

Bind values to evidence

Require every asserted value to carry a character span into the immutable source revision. A valid span must be within bounds and the captured text must support the value; the latter needs semantic review when normalization changes form. For example, the value “47 minutes” may be normalized to a duration, but the source phrase and unit must remain available. An empty field means missing, not false. A model can emit perfect JSON with a fabricated value, so syntax validation cannot replace evidence checks. Span annotation supplies the offset convention.

Handle repeated and conflicting mentions

A document may list a proposed window and later a revised window. Preserve both candidates with section, author and revision context instead of taking the first match. Decide which mention is active through an explicit document policy; unresolved conflicts should block downstream scheduling. Nested tables and OCR reading order can separate a label from its value, so offsets should refer to the extracted source representation and its revision, not a casually normalized string. Table scope covers that case.

Evaluate per field and per document

Report exact-span match, normalized-value correctness, missing-field precision, conflict detection and document-level completeness. A record with four correct fields and one invented approval is not releasable. Slice by document template, OCR quality and field frequency. Keep a blind test set with unseen services and changed layouts; otherwise the extractor may memorize a template. The repair lesson shows how to hold invalid output, and the project tests a release boundary.

Implementation

python
def validate_evidence_fields(source_text, extracted, allowed_fields):
    errors = []
    for field_name, field_value in extracted.items():
        if field_name not in allowed_fields:
            errors.append((field_name, "unknown-field"))
            continue
        start, end = field_value["span"]
        if not 0 <= start < end <= len(source_text):
            errors.append((field_name, "bad-span"))
        elif source_text[start:end] != field_value["evidence"]:
            errors.append((field_name, "span-mismatch"))
    return errors

request_text = "Service gateway-west. Window: 47 minutes."
start = request_text.index("gateway-west")
record = {"service": {"value": "gateway-west", "span":
                      (start, start + len("gateway-west")),
                      "evidence": "gateway-west"}}
assert validate_evidence_fields(request_text, record, {"service"}) == []
assert validate_evidence_fields(request_text,
                                {"approval": record["service"]},
                                {"service"}) == [("approval", "unknown-field")]

Performance and operating cost

For f fields, the loop is O(f) plus the cost of slicing and comparing evidence text; total copied character work is O(e) across e evidence characters. The example checks span integrity and a field allowlist, not whether the quoted words semantically justify the normalized value. That check remains a separate review step.

Common Mistakes

  • Treating valid JSON as proof that each value appears in the document.
  • Converting a missing field into a false approval.
  • Losing offsets when text is cleaned or OCR is rerun.
  • Taking the first timestamp when a later revision replaces it.

Read next

ai-data
natural-language-processing
Storage details