A written token does not always have one obvious spoken form. A speech prompt should keep the displayed text and its approved reading separate. Provide pronunciations for names, expansions for abbreviations, and rules for dates, units, and identifiers. If the synthesis engine supports a lexicon or speech markup, validate the syntax and language setting for that engine; otherwise use an approved spoken script. Do not assume a model will say '47 cm' or 'XQ-74' in the way the editor expects. Check the rendered audio, not only the prompt text. A pronunciation guide cannot repair a wrong underlying fact, so compare numbers with the approved record before rendering.
Speech pronunciation prompts: resolve names and numbers explicitly
Operational case
The museum calls its invented exhibit 'Orin' and approves the spoken reading 'OR-in.' A generator instead says 'oh-REEN' in one clip. In the same script, it reads the inventory code XQ-74 as a quantity and says 'seventy-four' without the letters. The public narration does not need that internal code, so the editor removes it rather than overloading the audience with a technical identifier. The 47-centimeter height and 13 brass markers do matter; their spoken forms are written in the recording brief. A reviewer listens for both values and the approved name, then checks that the visible exhibit label uses the same name.
Display: Orin Dial | 47 cm | 13 brass markers.
Approved speech: 'OR-in Dial; forty-seven centimeters; thirteen brass markers.'
Internal code XQ-74: omit from public narration.
Review: name, measurements, and label agree in rendered media.Performance and operating cost
For K special terms and N narration segments, a pronunciation checklist requires O(KN) naive listening checks; indexing each term's timecode reduces repeat inspection. A lexicon can make rendering consistent, but each voice and language configuration should be tested because support differs across engines. Treat number-to-speech conversion as a formatting problem with a reviewed output, not a creative writing task. The additional lexicon and approval work is small compared with correcting a released audio guide in several languages and replacing its captions and transcript.
Common Mistakes
- Do not assume a spelling gives the approved pronunciation.
- Do not read an internal identifier aloud merely because it appears in source data.
- Do not let a pronunciation hint change the measured value.
Connected lessons
- Prompt engineering applications
- Prompt Engineering
- Locale-aware prompts: keep values typed until rendering
- Translation prompts: preserve terms and machine placeholders
- Speech generation prompts: lock words before directing delivery
- Generated speech review: check the actual voice asset
- Video generation prompts: specify one shot's motion and evidence
- Generated video release: inspect continuity, claims, and alternatives
- Project: release an audio and video guide for a fictional exhibit
- Generated speech and video prompt decisions
