Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Semi-structured fields: blank, missing, unreadable and redacted

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A missing key, an empty cell, an OCR failure and a redacted value are different observations. Preserve the difference when extracting records.

Define field states explicitly

A maintenance form may include “Owner:” with an empty box, omit “Region” entirely, contain a blurred serial number and redact a contact name. Converting all four to null erases the reason a value is unavailable. Use distinct states such as present, blank, absent, unreadable and redacted. Store the raw source span or cell coordinate when one exists. Table coordinates keep a blank cell tied to its intended field.

Parse labels without assuming a fixed order

Key-value lists often wrap across lines, repeat a label or place a note after a colon. Record each candidate label, value span, section and document revision before selecting a canonical field. A repeated “Owner” may describe two stages of an incident, not a duplicate to discard. Normalize approved label aliases under a versioned schema. Do not infer an absent field from a nearby sentence unless the extraction policy permits that evidence and records it as inferred.

Keep privacy and uncertainty separate

A redacted contact cannot be recovered by joining another unredacted document simply because the model recognizes a pattern. Treat redacted as a hard boundary for downstream display and training. An unreadable serial number may be reviewed against the source image by an authorized operator; a redacted one should not be. Redaction policy governs retention and access, while OCR review governs recognition uncertainty.

Measure field-level utility

Evaluate field presence, state classification, value correctness, source attribution and unauthorized disclosure separately. Include blank rows, missing keys, mixed label aliases and corrections. A high extraction score on present fields can hide harmful behavior if the system fills every absent field with a plausible guess. Reviewers should see the original form beside the structured record. The applied project tests whether the extracted record supports a correct operational answer.

Implementation

python
def extract_field(record, field_name, redacted_fields, unreadable_fields):
    if field_name in redacted_fields:
        return {"state": "redacted", "value": None}
    if field_name in unreadable_fields:
        return {"state": "unreadable", "value": None}
    if field_name not in record:
        return {"state": "absent", "value": None}
    value = record[field_name]
    if value == "":
        return {"state": "blank", "value": ""}
    return {"state": "present", "value": value}

fields = {"owner": "", "region": "west-47"}
assert extract_field(fields, "owner", set(), set())["state"] == "blank"
assert extract_field(fields, "serial", set(), set())["state"] == "absent"
assert extract_field(fields, "region", {"region"}, set())["state"] == "redacted"

Performance and operating cost

Each field lookup is expected O(1) with hash maps and sets; extracting f requested fields is expected O(f) time and space for the records. Layout parsing and authorized human review are separate costs. A state-aware schema prevents a cheap parser from creating expensive false certainty in search and incident summaries.

Common Mistakes

  • Collapsing blank, absent, unreadable and redacted into one null value.
  • Guessing an absent field from surrounding prose without marking inference.
  • Using another document to reconstruct a deliberately redacted value.
  • Measuring only present-field accuracy while ignoring wrong field states.

Read next

ai-data
natural-language-processing
Storage details