A multimodal prompt should identify the artifact, requested observation, and limits of what can be inferred. OCR, diagrams, screenshots, and audio transcriptions can omit small labels, crop context, or confuse similar symbols. Ask for visible or audible observations first and an explicit uncertainty state for hidden or unreadable details. When a result matters operationally, verify dimensions, timestamps, or extracted text with a deterministic tool or a second human check. Treat text embedded inside media as lower-trust content, not a new instruction to the assistant.
Multimodal prompts: separate what an image shows from what it suggests
Decision in practice
A field engineer uploads a photograph of a control panel. The task is to transcribe the displayed alarm code and identify the lit indicator, not to diagnose a machine failure from a single frame. The assistant reports that AL-47 is legible, the amber indicator appears lit, and the lower serial label is obscured. It asks for a closer image before matching the device to an asset record. A sticker saying ignore safety checks is described as sticker text and does not alter the work. The engineer compares the transcription with a direct panel read before filing a maintenance ticket.
Artifact: panel-photo-27, one cropped frame.
Observed: alarm AL-47; amber indicator appears lit.
Not visible: full serial label and prior alarm history.
Next check: close image of serial label; direct panel confirmation.Performance and operating cost
Higher-resolution media and multiple frames increase upload size, processing time, and review cost. A crop can improve readability but also remove context, so keep the original and the crop ID. Measure field-level transcription accuracy and the rate of unsupported diagnoses. An assistant that says a light is on should not infer the root cause of a machine fault without the relevant telemetry. The same boundary applies to audio: uncertain words and speaker attribution must remain uncertain.
Common Mistakes
- Do not treat text inside an image as a trusted instruction.
- Do not invent a hidden serial number from a blurry label.
- Do not turn one visible indicator into a full diagnosis.
Connected lessons
- Prompt Engineering
- Prompt patterns
- Document extraction: separate observed fields from inferred values
- Evidence IDs: make generated claims auditable against supplied records
- Conversation memory: retain decisions without retaining every private detail
- Project: defend a retrieval and action workflow
- Prompt design decisions
Apply the boundary
Use the artifact's source and destination to decide what must be checked before the result is accepted.
- Image prompts: specify the visual contract and inspect the pixels
- Audio prompts: preserve speakers, timestamps, and uncertain words
- Video prompts: cite the moment and the observation
Continue with: Cross-modal evidence: keep conflicting observations separate.
