Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Speech input text: expand numbers, units and protected tokens

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Written text is not a complete speech script. Expand ambiguous numbers and units under a locale policy while protecting service identifiers.

Separate display text from spoken form

A dashboard may display “gateway-west: 47 ms,” but a spoken alert should name the service and say “forty-seven milliseconds” under a declared language policy. Store both the original text and the reviewed spoken form; do not overwrite the written value. A bare “47” could be an incident ID, count, time or version. Its surrounding field type determines the expansion. Quantity contracts preserve value and unit before any speech rendering.

Protect names and commands

Service names, acronyms, paths and commands can sound misleading when an engine guesses pronunciation. Keep protected spans with typed roles and locale-specific readings approved by owners. A command shown in a runbook may need character-by-character reading or may be omitted from speech entirely if reading it would be unsafe. Abbreviation scope helps distinguish a product acronym from an ordinary word.

Treat normalization as a versioned transform

The same text can require different spoken forms by locale, voice or product policy. Record normalization rules, lexicon version and source text revision so generated audio can be reproduced and corrected. If a value or term is unknown, route it to a pronunciation review queue rather than silently guessing. Keep offsets from original text to speech segments for captions and incident audit. Source offsets prevent a cleaned string from losing its link to the displayed record.

Evaluate meaning and intelligibility

Test numbers, decimals, units, dates, identifiers, abbreviations and mixed-language terms. Human listeners should verify the spoken value and service name, not merely rate naturalness. A pleasant voice can still read “ms” as an unrelated word or reverse a threshold. The lexicon lesson controls domain words; the project gates critical alert audio.

Implementation

python
def spoken_metric(metric, locale):
    readings = {("en", "ms"): "milliseconds",
                ("en", "min"): "minutes"}
    if not isinstance(metric["value"], int):
        return {"state": "review", "reason": "unsupported-value"}
    unit = readings.get((locale, metric["unit"]))
    if unit is None:
        return {"state": "review", "reason": "unknown-unit"}
    return {"state": "ready", "spoken_text":
            f"{metric['value']} {unit}", "display_value":
            f"{metric['value']} {metric['unit']}"}

assert spoken_metric({"value": 47, "unit": "ms"}, "en") == {
    "state": "ready", "spoken_text": "47 milliseconds",
    "display_value": "47 ms"}
assert spoken_metric({"value": 82, "unit": "ms"}, "hi")["state"] == "review"

Performance and operating cost

A fixed lookup and integer check are expected O(1) time and space, excluding output-string length. This example expands a unit but leaves number-to-words conversion to a locale-aware speech engine or reviewed formatter. Never assume its default pronunciation is correct for critical values or identifiers.

Common Mistakes

  • Replacing display text with spoken text and losing source fidelity.
  • Reading every number as a count without field context.
  • Guessing a service pronunciation from spelling alone.
  • Reusing audio after the underlying threshold or lexicon changed.

Read next

ai-data
natural-language-processing
Storage details