An image verification prompt asks for an observation tied to a particular image region, then records what the image cannot establish. A model may describe a dent or read a label, but the crop, resolution, angle, and lighting determine what is actually visible. Use stable image IDs and region coordinates or crop IDs, and inspect the source image when a claim affects a decision. Do not infer serial numbers, material defects, or causation from a vague global caption. A region annotation is evidence location, not a guarantee that the interpretation is correct.
Image prompts: verify claims against named regions and crops
Decision in practice
Photo PH-309 shows three stacked cartons beside a pallet. A broad caption says the top carton is damaged. The visible tear is actually on the lower right carton, while the top carton has a shadow along its seam. The team asks the assistant to identify the affected carton by position, crop ID, and visible feature. A reviewer zooms into crop CR-8 and confirms a torn edge but cannot read the pallet label. The output records the tear and the unreadable label separately. It does not declare that the shipment was damaged in transit; the photo alone lacks timing and chain of custody.
Image PH-309; region CR-8: lower-right carton, torn outer edge.
Image PH-309; region CR-9: pallet label unreadable.
Observation: visible tear; object identity: provisional.
Unsupported: damage time, responsible party, hidden contents.
Decision: request a label close-up before matching shipment ID.Performance and operating cost
A full-image pass plus K targeted crops takes O(K) additional model or review steps and can increase image-token and storage costs. Crop only where the acceptance check needs detail; a crop that removes neighboring objects may destroy positional context, so retain its source coordinates. Review workload rises with ambiguous images, but false certainty is costlier when a photo drives a claim. Measure localization errors and missed unreadable labels separately. If the image is too small, request a better capture instead of enlarging an invented detail.
Common Mistakes
- Do not claim a precise object identity from a wide image with unreadable labels.
- Do not let a tight crop erase the context needed to locate an object.
- Do not infer cause or timing from a single still image.
Connected lessons
- Prompt engineering applications
- Prompt Engineering
- Image prompts: specify the visual contract and inspect the pixels
- Multimodal prompts: separate what an image shows from what it suggests
- Selective answers: measure when to abstain
- Document prompts: anchor each field to a page and resolve conflicts
- Audio prompts: mark overlapping speech before assigning speakers
- Video prompts: disclose sampling gaps around short events
- Cross-modal evidence: keep conflicting observations separate
- Project: reconcile a warehouse evidence packet
- Multimodal prompt evidence decisions
Continue with: Chart prompts: inspect axes, units, and missing series first.
Continue with: Image alternatives: prompt for purpose, then check the image.
