A domain pronunciation rule needs an owner, locale, source revision and audible review before it is used for operational speech.
Pronunciation lexicons: locale, domain terms and review
Model a term as more than spelling
The string “SRE” may be read as separate letters, while a service name may have an approved spoken alias. Store token or phrase, language, pronunciation representation, context scope, owner and lexicon revision. A single global replacement may corrupt ordinary words that share the same letters. Use longest approved phrase matches under explicit boundaries, then send unmatched technical terms for review. Text normalization decides the unit and identifier roles before this pronunciation lookup.
Version changes against audio
When a team renames gateway-west, old audio may still use the former name. Keep the source text revision, lexicon revision and synthesis engine version on each audio artifact. Invalidate or re-record affected alerts when any of these change. Reviewers should hear the full phrase, because a token-level pronunciation can be correct in isolation and confusing in context. Revision auditing uses a similar dependency idea for text search.
Handle mixed language carefully
A Hindi sentence can contain an English service ID. Choose pronunciation per span, not per whole message, and keep the original identifier visible in captions. If the language switch is uncertain, ask a reviewer or use a safe explicit spelling policy. A lexicon entry is not proof that the whole sentence has correct accent, timing or meaning. Code-switching analysis informs the segment policy.
Audit by listening
Measure term pronunciation accuracy, value accuracy, listener comprehension and mismatch between caption and speech. Include new services, abbreviations with collisions and low-volume locales. A word error rate computed by an automatic transcriber can help triage but cannot replace a targeted listening check for a critical threshold. The operational alert project keeps an approved audio manifest and a rollback path.
Implementation
def approved_pronunciation(term, locale, lexicon):
entry = lexicon.get((locale, term))
if entry is None or entry["review_state"] != "approved":
return {"state": "review", "term": term, "locale": locale}
return {"state": "ready", "term": term, "locale": locale,
"spoken_form": entry["spoken_form"],
"lexicon_revision": entry["revision"]}
lexicon = {("en", "SRE"): {"spoken_form": "S R E",
"review_state": "approved", "revision": "lex-r8"},
("en", "gateway-west"): {"spoken_form": "gateway west",
"review_state": "pending", "revision": "lex-r8"}}
assert approved_pronunciation("SRE", "en", lexicon)["spoken_form"] == "S R E"
assert approved_pronunciation("gateway-west", "en", lexicon)["state"] == "review"
Performance and operating cost
One dictionary lookup is expected O(1) time and space for a term. Scanning an utterance for many phrases requires tokenization and a match index; a naive scan over t terms and n tokens can cost O(t × n). A reviewed lexicon improves named terms but does not validate the speech engine’s complete output.
Common Mistakes
- Applying a global string replacement without token boundaries.
- Using a pending pronunciation in a critical alert.
- Forgetting to regenerate audio after a lexicon revision.
- Scoring naturalness while ignoring whether the service name was understood.
Read next
- Speech input text: expand numbers, units and protected tokens
- Project: release a spoken operational alert with reviewed values
- Abbreviations: definition scope, collisions and unknown forms
- Script profiles and code-switching boundaries in text intake
- Speech transcripts: segments, recognition errors and entity risk
