Extract a document into typed fields whose values can be traced to exact source spans; a well-formed object alone is not evidence.
Schema-bound extraction: typed fields and exact source spans
Define fields before decoding
A change request may mention a service, requested window and approval status in different paragraphs. Define each field’s meaning, type, cardinality and allowed null state before a model reads the text. A “window” can mean start time, end time or a duration, so one loose string field is inadequate. Keep schema version with every extraction. Time anchors and zones are needed before a local time can become a scheduled instant.
Bind values to evidence
Require every asserted value to carry a character span into the immutable source revision. A valid span must be within bounds and the captured text must support the value; the latter needs semantic review when normalization changes form. For example, the value “47 minutes” may be normalized to a duration, but the source phrase and unit must remain available. An empty field means missing, not false. A model can emit perfect JSON with a fabricated value, so syntax validation cannot replace evidence checks. Span annotation supplies the offset convention.
Handle repeated and conflicting mentions
A document may list a proposed window and later a revised window. Preserve both candidates with section, author and revision context instead of taking the first match. Decide which mention is active through an explicit document policy; unresolved conflicts should block downstream scheduling. Nested tables and OCR reading order can separate a label from its value, so offsets should refer to the extracted source representation and its revision, not a casually normalized string. Table scope covers that case.
Evaluate per field and per document
Report exact-span match, normalized-value correctness, missing-field precision, conflict detection and document-level completeness. A record with four correct fields and one invented approval is not releasable. Slice by document template, OCR quality and field frequency. Keep a blind test set with unseen services and changed layouts; otherwise the extractor may memorize a template. The repair lesson shows how to hold invalid output, and the project tests a release boundary.
Implementation
def validate_evidence_fields(source_text, extracted, allowed_fields):
errors = []
for field_name, field_value in extracted.items():
if field_name not in allowed_fields:
errors.append((field_name, "unknown-field"))
continue
start, end = field_value["span"]
if not 0 <= start < end <= len(source_text):
errors.append((field_name, "bad-span"))
elif source_text[start:end] != field_value["evidence"]:
errors.append((field_name, "span-mismatch"))
return errors
request_text = "Service gateway-west. Window: 47 minutes."
start = request_text.index("gateway-west")
record = {"service": {"value": "gateway-west", "span":
(start, start + len("gateway-west")),
"evidence": "gateway-west"}}
assert validate_evidence_fields(request_text, record, {"service"}) == []
assert validate_evidence_fields(request_text,
{"approval": record["service"]},
{"service"}) == [("approval", "unknown-field")]
Performance and operating cost
For f fields, the loop is O(f) plus the cost of slicing and comparing evidence text; total copied character work is O(e) across e evidence characters. The example checks span integrity and a field allowlist, not whether the quoted words semantically justify the normalized value. That check remains a separate review step.
Common Mistakes
- Treating valid JSON as proof that each value appears in the document.
- Converting a missing field into a false approval.
- Losing offsets when text is cleaned or OCR is rerun.
- Taking the first timestamp when a later revision replaces it.
