A missing key, an empty cell, an OCR failure and a redacted value are different observations. Preserve the difference when extracting records.
Semi-structured fields: blank, missing, unreadable and redacted
Define field states explicitly
A maintenance form may include “Owner:” with an empty box, omit “Region” entirely, contain a blurred serial number and redact a contact name. Converting all four to null erases the reason a value is unavailable. Use distinct states such as present, blank, absent, unreadable and redacted. Store the raw source span or cell coordinate when one exists. Table coordinates keep a blank cell tied to its intended field.
Parse labels without assuming a fixed order
Key-value lists often wrap across lines, repeat a label or place a note after a colon. Record each candidate label, value span, section and document revision before selecting a canonical field. A repeated “Owner” may describe two stages of an incident, not a duplicate to discard. Normalize approved label aliases under a versioned schema. Do not infer an absent field from a nearby sentence unless the extraction policy permits that evidence and records it as inferred.
Keep privacy and uncertainty separate
A redacted contact cannot be recovered by joining another unredacted document simply because the model recognizes a pattern. Treat redacted as a hard boundary for downstream display and training. An unreadable serial number may be reviewed against the source image by an authorized operator; a redacted one should not be. Redaction policy governs retention and access, while OCR review governs recognition uncertainty.
Measure field-level utility
Evaluate field presence, state classification, value correctness, source attribution and unauthorized disclosure separately. Include blank rows, missing keys, mixed label aliases and corrections. A high extraction score on present fields can hide harmful behavior if the system fills every absent field with a plausible guess. Reviewers should see the original form beside the structured record. The applied project tests whether the extracted record supports a correct operational answer.
Implementation
def extract_field(record, field_name, redacted_fields, unreadable_fields):
if field_name in redacted_fields:
return {"state": "redacted", "value": None}
if field_name in unreadable_fields:
return {"state": "unreadable", "value": None}
if field_name not in record:
return {"state": "absent", "value": None}
value = record[field_name]
if value == "":
return {"state": "blank", "value": ""}
return {"state": "present", "value": value}
fields = {"owner": "", "region": "west-47"}
assert extract_field(fields, "owner", set(), set())["state"] == "blank"
assert extract_field(fields, "serial", set(), set())["state"] == "absent"
assert extract_field(fields, "region", {"region"}, set())["state"] == "redacted"
Performance and operating cost
Each field lookup is expected O(1) with hash maps and sets; extracting f requested fields is expected O(f) time and space for the records. Layout parsing and authorized human review are separate costs. A state-aware schema prevents a cheap parser from creating expensive false certainty in search and incident summaries.
Common Mistakes
- Collapsing blank, absent, unreadable and redacted into one null value.
- Guessing an absent field from surrounding prose without marking inference.
- Using another document to reconstruct a deliberately redacted value.
- Measuring only present-field accuracy while ignoring wrong field states.
