Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Document extraction: separate observed fields from inferred values

Last updated: 5 Oct 20268 min read
tutorial
IntermediateBy AITrove Editorial

A document extractor needs a schema, field-level evidence positions, and explicit null handling. OCR or parsing can confuse a character, merge table rows, or miss a page. The prompt should distinguish an observed value from an inferred one and avoid computing a total from uncertain components. A downstream validator checks type, cross-field constraints, and whether the source span exists. Some fields can be parsed deterministically; use a model for layout variation and ambiguous language where those rules fail. Keep the original document in a controlled store for review.

Decision in practice

An accounts-payable service receives 118 invoices. Each result contains invoice_id, supplier_id, currency, subtotal, tax, total, and page-region keys. On one scanned invoice the tax digits are blurred. The model returns tax as unknown with a page region rather than guessing a value from the total. A validator confirms subtotal plus tax equals total only when all three numbers are observed. Duplicate invoice IDs with different suppliers enter review, not automatic payment. The team samples low-confidence scans and measures field-level error rates separately from full-document success.

Output
invoice_id: INV-4271
supplier_id: SUP-82
subtotal: 684.00 observed on page 2
tax: unknown; blurred region page 2, box 17
total: 718.20 observed on page 2
action: review; no computed tax substituted for observation

Performance and operating cost

Extraction cost scales with document count and page length; OCR and model calls can dominate latency. A field-level review queue concentrates human effort on ambiguous values instead of rereading every clean document. Full-document accuracy can be much lower than per-field accuracy when many fields are required, so report both. Do not use a passing arithmetic check as proof the underlying invoice is legitimate. The payment effect belongs behind supplier, duplicate, and authorization checks outside the prompt.

Common Mistakes

  • Do not present an inferred number as an observed field.
  • Do not use a valid sum as proof the invoice is authentic.
  • Do not omit field-level evidence positions for ambiguous scans.

Connected lessons

Next decision

Check whether the right facts reached the workflow and whether the result is safe at its destination.

Continue with: Document prompts: anchor each field to a page and resolve conflicts.

Continue with: Invoice prompts: pin provenance and vendor identity.

prompt engineering
tutorial
Storage details