A video prompt should be a shot contract, not a pile of atmosphere words. Identify the subject and approved reference asset, the action over time, camera motion, framing, duration target, and facts that must remain visible. A reference image may establish composition, but the output can still invent objects or change labels. For explanatory footage, do not rely on generated in-frame text to carry exact instructions or numbers; add verified text in post-production where possible. Keep a shot list with start and end states so the editor can review temporal continuity. A prompt describes intent, not a guarantee that the rendered frames obey geometry or timing.
Video generation prompts: specify one shot’s motion and evidence
Operational case
The museum wants an eight-second clip of the fictional Orin Dial. Its approved photo shows 13 brass markers around a fixed rim. The brief asks for a slow quarter-turn of the camera around the model while the model itself remains still. A draft clip animates the markers as moving clock hands and adds a fourteenth marker. The editor rejects it: both changes contradict the exhibit record. A revised shot uses a locked reference, modest camera movement, no generated on-screen label, and a separate verified overlay. The reviewer inspects the beginning, midpoint, and end, then scans the whole clip for transient extra parts.
Subject: approved Orin Dial asset; 13 fixed brass markers.
Action: model remains still; camera makes slow quarter-turn.
Target: eight seconds; label added later from verified text.
Reject: moving markers, extra parts, unstable object identity.Performance and operating cost
For S shots and F inspected frames per shot, sampled visual review is O(SF), while a full-duration watch costs time proportional to total video length. Sampling can miss a one-frame error, so inspect the whole short clip when exact object details matter. Long or elaborate prompts increase editing and review complexity without making a renderer deterministic. Keep reference asset version, shot prompt, generator settings, clip hash, and rejected reasons together. If exact marker count cannot be verified reliably in generated motion, use conventional footage or graphics for that shot.
Common Mistakes
- Do not treat a reference image as a guarantee of object fidelity.
- Do not put exact public instructions only in generated in-frame text.
- Do not check a temporal clip from a single still frame.
Connected lessons
- Prompt engineering applications
- Prompt Engineering
- Image prompts: specify the visual contract and inspect the pixels
- Video prompts: disclose sampling gaps around short events
- Image prompts: verify claims against named regions and crops
- Speech generation prompts: lock words before directing delivery
- Speech pronunciation prompts: resolve names and numbers explicitly
- Generated speech review: check the actual voice asset
- Generated video release: inspect continuity, claims, and alternatives
- Project: release an audio and video guide for a fictional exhibit
- Generated speech and video prompt decisions
