Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Quantities in text: value, unit, range and source span

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A number is not a metric until its unit, entity, time window and source are known. Parse those fields before comparing claims.

Bind the value to what it measures

“47 retries in 82 minutes” contains a count and a duration. Store each value with its unit, source span, associated service and observation window. Do not treat the duration as a second retry count. “About 47” is approximate; “under 47” is an upper bound; “47–82” is a range. These need distinct comparison rules. Time anchors are required when a quantity refers to a relative window.

Preserve locale and written form

A decimal separator, grouping mark or percent sign can be misread across locales. Keep the exact source string and declared locale before parsing. Avoid binary floating-point for money or a threshold that demands decimal exactness. A unit such as ms versus s changes scale, while “requests” versus “failed requests” changes the measured population. A parser should abstain on unknown units rather than silently assuming a familiar one.

Normalize only under a declared policy

Convert between compatible units with an explicit factor and version. Do not compare bytes with requests, or percentages with counts lacking denominators. If one source reports an interval and another a point estimate, preserve that distinction through aggregation. Table headers may supply the unit when the cell holds only a bare number; record the header as part of its evidence.

Audit extraction and interpretation

Score numeric span, unit, comparator, entity, window and derived value separately. Include OCR confusions, negatives, ranges, approximations, percentages and units that look alike. A system can read “47” correctly but attach it to the wrong service or time period. Reviewers should see the original phrase and any conversion step. The numeric claim project tests the full path from text to answer.

Implementation

python
from decimal import Decimal

UNIT_TO_SECONDS = {"s": Decimal("1"), "min": Decimal("60")}

def duration_seconds(raw_value, unit, source_span):
    if unit not in UNIT_TO_SECONDS or not source_span:
        return {"state": "review", "reason": "unit-or-source-missing"}
    value = Decimal(raw_value)
    if value < 0:
        return {"state": "review", "reason": "negative-duration"}
    return {"state": "normalized", "seconds": value * UNIT_TO_SECONDS[unit],
            "original": raw_value, "unit": unit, "source_span": source_span}

assert duration_seconds("82", "min", "82 minutes")["seconds"] == Decimal("4920")
assert duration_seconds("47", "requests", "47 requests")["state"] == "review"

Performance and operating cost

A dictionary unit lookup and decimal conversion of d digits cost at least O(d) time and space. Parsing n characters of text is at least O(n). Unit normalization is cheap relative to correcting wrong entity or time-window links, so measure those errors separately. The example handles two duration units, not arbitrary physical dimensions.

Common Mistakes

  • Comparing counts and percentages without a denominator.
  • Assuming a bare number has the unit used by another document.
  • Discarding “about,” “under” or a range marker.
  • Using binary floating-point when exact decimal thresholds matter.

Read next

Continue the workflow: Speech input text: expand numbers, units and protected tokens.

Continue the workflow: Spatial language: scope, uncertainty and map resolution.

ai-data
natural-language-processing
Storage details