Build a change-request intake that captures source spans, detects revisions and blocks scheduling until required fields are verified.
Project: extract change requests with field-level evidence
Build the corpus
Prepare fictional change requests for gateway-west and archive-east with a service identifier, requested window, risk note and approval wording. Include a draft followed by a correction, a negated approval and an OCR copy that swaps two digits. Annotate exact spans in immutable document revisions, normalized field values and field states. Group versions of the same request in one evaluation split. Field evidence provides the annotation contract.
Implement the pipeline
Parse document structure, propose values under a fixed schema, validate types and spans, then compare repeated mentions for conflict. Keep both the source phrase and the normalized value. A local timestamp needs a zone and date; a bare “next Friday” needs a reference date and may remain unresolved. Pass malformed outputs through one bounded repair attempt only when the validator can name a fix. Repair states distinguish that from semantic disagreement.
Gate downstream scheduling
Release a record to human scheduling review only after the source revision, service, window and approval evidence are present and not conflicting. A negative statement such as “approval pending” must not be normalized to approved. The code below checks release prerequisites; the actual approval interpretation and reviewer signature must come from the workflow’s trusted policy. Store a decision ledger that ties each field to its source and reviewer.
Evaluate end to end
Report per-field value accuracy, source-span fidelity, false approval rate, conflict recall, held-record rate and human correction time. Score the final scheduled record, not just the extractor output. A single false approval should be visible even when most fields are correct. Re-run the corpus after schema or OCR changes and invalidate downstream records whose supporting span changed.
Implementation
def admit_change_request(record):
required = ("source_revision", "service_id", "window_utc",
"approval_evidence", "reviewer_id")
missing = [name for name in required if not record.get(name)]
if missing:
return {"state": "hold", "reason": "missing", "fields": missing}
if record.get("field_conflict"):
return {"state": "hold", "reason": "conflict"}
if record["approval_evidence"]["state"] != "approved":
return {"state": "hold", "reason": "approval-not-verified"}
return {"state": "review-ready", "service_id": record["service_id"]}
request = {"source_revision": "change-r7", "service_id": "gateway-west",
"window_utc": "2026-10-12T03:47:00Z",
"approval_evidence": {"state": "pending"},
"reviewer_id": "reviewer-82", "field_conflict": False}
assert admit_change_request(request)["reason"] == "approval-not-verified"
assert admit_change_request({**request, "approval_evidence":
{"state": "approved"}})["state"] == "review-ready"
Performance and operating cost
The fixed-field admission gate takes O(1) time and space. Extracting and reviewing a document scales with its text length, number of candidate spans and human exceptions. A review-ready state is not an instruction to schedule work; it means a trusted reviewer can inspect the evidence-bound record.
Common Mistakes
- Calling “approval pending” an approval.
- Mixing fields from two revisions into one record.
- Dropping the source spans after normalization.
- Evaluating only individual fields while ignoring unsafe final records.
