A transcript is a derived reading of audio. Keep timing and uncertainty so an incorrect order ID cannot become a trusted fact.
Speech transcripts: segments, recognition errors and entity risk
Represent the transcript as aligned segments
Store recording revision, segment start and end time, recognized words, language hypothesis and recognizer version. A final display paragraph is a derived view. Do not assume punctuation or sentence boundaries existed in the speech. Keep the audio span addressable for an authorized reviewer, and avoid publishing restricted recordings through a transcript link. Text offsets still matter, but word-to-audio timing is a separate coordinate system.
Treat entities as a separate failure class
Word error rate can improve while the model still confuses ZX-47 with ZX-74 or turns “not approved” into “approved.” Measure exact recognition of identifiers, amounts, names and negation in addition to overall word errors. Verify critical values through an authorized system or an explicit human replay of the audio. A language model may produce a plausible product name that was never spoken. Entity extraction should therefore consume a transcript with provenance and uncertainty.
Keep corrections traceable
A reviewer may correct a misheard word, but the original recognition result must remain available with the audio interval and correction record. Downstream offsets and extracted entities must be recomputed from the corrected transcript revision. Mark an unresolved segment rather than filling it from neighboring context. If consent or retention policy removes audio, revoke derived transcript and index records according to that policy.
Evaluate the intended workflow
Audit clean and noisy calls, accents, code-switching, cross-talk, long silences and rare product names. Report word errors, entity errors, time alignment errors, review time and false actions. Compare models on the same held-out recordings, grouped by conversation. Speaker attribution adds another uncertainty source; the project places both in a review queue.
Implementation
def validate_transcript_segment(segment, recording_duration_ms):
start = segment["start_ms"]
end = segment["end_ms"]
if not 0 <= start < end <= recording_duration_ms:
raise ValueError("segment lies outside recording")
if not segment["recording_revision"]:
raise ValueError("recording revision is required")
return {"start_ms": start, "end_ms": end,
"text": segment["text"],
"recording_revision": segment["recording_revision"]}
spoken = {"start_ms": 4700, "end_ms": 8200,
"text": "order ZX-47", "recording_revision": "call-r7"}
assert validate_transcript_segment(spoken, 19000)["text"] == "order ZX-47"
Performance and operating cost
Segment validation is O(1) time and space apart from copying its text. Recognition cost grows with audio duration, while storing word timing adds metadata. Human replay is expensive, so prioritize critical entities and low-confidence segments. A faster recognizer is not an improvement if it produces more wrong identifiers in the workflow.
Common Mistakes
- Treating punctuation emitted by the recognizer as spoken evidence.
- Using overall word error rate as a substitute for identifier accuracy.
- Applying corrected-text offsets to the original transcript.
- Keeping searchable transcript derivatives after audio deletion requires removal.
