Prepare an audio and video guide for the fictional Orin Dial exhibit. The approved exhibit record states that the model is 47 centimeters tall, has 13 fixed brass markers, and does not measure time. The name is spoken 'OR-in.' The deliverable is a review packet for a short narration and two eight-second video shots, not an instruction to publish an uninspected generation. Keep the factual script, delivery direction, voice choice, reference photograph, shot plan, captions, transcript, and final export as distinct versioned artifacts.
Project: release an audio and video guide for a fictional exhibit
Lock speech and image facts
Write a script that preserves all three approved facts. Record the spoken forms of the name, measurement, and marker count. Exclude internal inventory code XQ-74 from the public narration. Choose a voice asset approved for this project and request a measured pace with one short sentence pause. After rendering, transcribe and listen to the file. Reject any clip that calls the display model a working clock, clips 'thirteen,' or alters 'does not measure time.' Check actual duration and pair the final file with its corrected transcript.
Build and inspect two shots
Shot A uses the approved photograph and a slow camera move around a stationary model. Shot B continues from the same side, with the same 13 fixed markers. Put exact public words in a verified overlay rather than relying on generated text inside the scene. Inspect each clip and the assembled sequence; a polished shot with an extra marker fails. Retiming the assembly requires another caption check. The descriptive transcript states that the camera moves and the model remains still. Compare the final export hash with the asset loaded by the public player, then log review decisions for the audio, video, captions, and transcript.
Approved: Orin Dial; 47 cm; 13 fixed brass markers; not a timekeeper.
Speech: exact script + approved OR-in reading; listen to rendered file.
Shot A/B: stationary model; continuing camera move; no extra marker.
Accessibility: captions from final audio; descriptive transcript.
Release: reviewed export hash matches player asset.Performance and operating cost
With R audio renders of duration D, listening is O(RD). For S shots and F sampled frames, sampled visual checking is O(SF), but a short exact-detail clip still needs full playback. Caption work scales with cue count and must repeat after a material edit. Store script, voice, reference photo, prompt, model, clip, transcript, captions, export, and player versions together; otherwise a correct review can attach to the wrong file. If generated video cannot hold the marker count, replace that shot with photographed footage or a controlled graphic instead of multiplying prompt adjectives.
Common Mistakes
- Do not mix factual script edits with voice-style revisions.
- Do not approve a video from one attractive still frame.
- Do not publish captions or a transcript tied to an older render.
Connected lessons
- Speech generation prompts: lock words before directing delivery
- Speech pronunciation prompts: resolve names and numbers explicitly
- Generated speech review: check the actual voice asset
- Video generation prompts: specify one shot's motion and evidence
- Generated video release: inspect continuity, claims, and alternatives
- Image prompts: specify the visual contract and inspect the pixels
- Caption and transcript prompts: keep timing, speakers, and sound evidence
- Prompt release artifacts: version the whole decision path
- Generated speech and video prompt decisions
