Extract a support runbook table, answer a service-specific question and show the exact cell and header that support the answer.
Project: answer from support tables with cell evidence
Define the user task
An operator asks for the retry ceiling for gateway-west in the west region. The answer must come from one current, authorized table row with the correct header path and document revision. The service may return “unresolved” if a cell is blank, unreadable or contradictory. It must not borrow a nearby worker-east value. Cell-coordinate evidence supplies the row, column and header identity needed for the response.
Build a hard source set
Collect HTML tables, copied spreadsheets, scanned runbook pages and key-value forms. Include merged headers, repeated page headers, row labels split across lines, blank cells, redactions and two regions for the same service. Have reviewers mark exact supporting cells and unavailable states. Keep every revision of a runbook in the same evaluation split. The field-state contract stops a blank cell from being read as zero.
Stage extraction and answering
Parse table structure before converting cells to searchable text. Apply document access and revision filters, then match the requested service and region. If more than one active row qualifies with different values, stage a conflict instead of selecting the first. Return the raw value, table ID, row, column, header path and source revision with the answer. Route low-confidence OCR cells to authorized review; do not repair them by language-model guess.
Release on evidence quality
Report correct answers, wrong-row answers, header errors, unresolved rate, access violations and reviewer time. Inspect any answer whose source table changed after indexing. Shadow the service on resolved support cases, then keep the previous parser and index generation for rollback. The final display should let an operator open the source cell directly. A fluent answer with no cell trace is not sufficient for this task.
Implementation
def answer_table_cell(rows, service, region, field):
matches = [row for row in rows
if row["service"] == service and row["region"] == region]
if len(matches) != 1:
return {"state": "review", "reason": "missing-or-duplicate-row"}
row = matches[0]
if field not in row or row[field] in {None, ""}:
return {"state": "unresolved", "reason": "unavailable-cell"}
return {"state": "answered", "value": row[field],
"table_id": row["table_id"], "row_index": row["row_index"],
"header": field, "source_revision": row["source_revision"]}
rows = [{"service": "gateway-west", "region": "west", "retry ceiling": "47",
"table_id": "limits-82", "row_index": 3, "source_revision": "r5"}]
assert answer_table_cell(rows, "gateway-west", "west", "retry ceiling")["value"] == "47"
Performance and operating cost
Scanning n rows is O(n) time and O(n) temporary space in this direct version. An indexed table can reduce lookup cost, but parser quality and reviewer work dominate when headers or row boundaries are uncertain. The code assumes access and revision filtering happened upstream; enforce them before calling it in a real service.
Common Mistakes
- Answering from a matching service in the wrong region.
- Treating a blank cell as a numeric zero.
- Choosing the first row when duplicates conflict.
- Showing a value without a cell coordinate and source revision.
Read next
- Table evidence: cell coordinates, headers and reading order
- Semi-structured fields: blank, missing, unreadable and redacted
- Project: review OCR fields in a document-intake queue
- Project: answer runbook questions with source spans and abstention
- Project: index runbook sections with revision-safe passages
