Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Project: review numeric claims in incident summaries

Last updated: 6 Oct 20265 min read
project
AdvancedBy AITrove Editorial

Extract operational quantities, calculate a supported rate and keep an incident summary from overstating uncertain or mismatched measurements.

Define the decision boundary

An incident draft says “gateway-west failures reached 20%.” The system must find failed and eligible request counts for the same environment and time window, calculate the rate, and show both source spans. It can stage the sentence for a reviewer; it cannot invent a denominator or suppress a conflicting update. The quantity contract keeps raw values, units and uncertainty attached to their sources.

Build a difficult case set

Collect status notes, runbook tables and OCR screenshots with counts, durations, ratios and ranges. Add two regions, a correction of an earlier count, a blank denominator and a percentage-point comparison. Group all updates from one incident in one evaluation split. Reviewers mark the exact inputs, formula, window and whether the proposed sentence is supported. An easy arithmetic set is insufficient if the hard problem is selecting the right measurements.

Stage calculations with lineage

Parse numeric candidates, bind them to service and window, then require compatible units. Reject a denominator that comes from another environment or revision. Calculate with decimal arithmetic and save the formula plus evidence IDs. If active sources conflict, return a review state with both claims. Calculation lineage lets a correction invalidate a previously derived percentage. The summary generator receives only reviewed numeric claims.

Measure the released claim

Report wrong input selection, unit errors, zero-denominator handling, incorrect rounding, unsupported trend wording and reviewer time. Check the final sentence against the calculation, because “rose to 20%” and “rose by 20%” are not equivalent. Shadow the process on resolved incidents before release. Keep source revisions and the previous calculation policy so a bad parser update can be rolled back without erasing earlier evidence.

Implementation

python
from decimal import Decimal

def stage_incident_rate(failed, eligible, same_scope, source_ids):
    if not same_scope or len(source_ids) != 2:
        return {"state": "review", "reason": "scope-or-evidence-mismatch"}
    failed_count = Decimal(failed)
    eligible_count = Decimal(eligible)
    if eligible_count <= 0 or not 0 <= failed_count <= eligible_count:
        return {"state": "review", "reason": "invalid-counts"}
    return {"state": "proposed", "rate": failed_count / eligible_count,
            "source_ids": tuple(source_ids)}

claim = stage_incident_rate("47", "235", True,
                            ["incident-47-failed", "incident-47-total"])
assert claim["rate"] == Decimal("0.2")
assert stage_incident_rate("47", "235", False, ["a", "b"])["state"] == "review"

Performance and operating cost

The scope gate is O(1); decimal work grows with the number of digits. Finding and validating source counts is the expensive NLP and review step. Measure the share of proposed claims accepted without correction, plus the rate of unsupported published statements. A fast calculation cannot compensate for a numerator drawn from the wrong incident.

Common Mistakes

  • Using correct arithmetic on counts from different scopes.
  • Converting an uncertain OCR number into a precise published percentage.
  • Confusing “rose to” with “rose by” in the final sentence.
  • Failing to invalidate a rate when one source note is corrected.

Read next

ai-data
natural-language-processing
Storage details