A multimodal example is a governed relationship among media items, labels and a target decision, not merely files sharing a directory.
Multimodal data contract: paired records, identity and consent
Define the pair
AI Trove could link a receipt image to its extracted text and a reviewer’s description. Give each asset an immutable media ID, parent receipt ID, capture time, source and transform version. A filename is not an identity contract: resizes and retries can produce several files for one receipt. Image-unit rules should determine whether the training unit is a receipt, page or crop.
Record alignment
The text must correspond to the same receipt image and, for multipage documents, the same page or region. Store alignment method and confidence; do not silently pair records by row order after a missing upload. A corrected OCR result should keep a link to the earlier version so a historical prediction can be reconstructed. The target label must stay distinct from the text used as model input.
Respect availability and permission
A user may permit a receipt image for a specific workflow but not indefinite training reuse. Track permitted purpose, retention deadline and deletion state for each asset and derived representation. A text transcript can contain the same sensitive fields as the image. Inference logging must obey those limits, including when a model emits intermediate descriptions.
Test a broken pair
Create image I-47 with OCR T-47, then insert a late text row T-48 for a different receipt. A positional join should fail the fixture; a keyed join should retain the correct pair. Add a deleted image whose cached text remains, and require the resulting training example to be excluded until the deletion policy is resolved.
Implementation
def paired_receipts(images, texts):
text_by_receipt = {row["receipt_id"]: row for row in texts}
return [(image, text_by_receipt[image["receipt_id"]])
for image in images
if image["receipt_id"] in text_by_receipt
and image["consent_ok"] and text_by_receipt[image["receipt_id"]]["consent_ok"]]Performance and operating cost
A keyed join costs O(I + T) expected time and O(T + P) memory for I images, T text rows and P accepted pairs. The simple map assumes one text version per receipt; production joins need version and page keys plus duplicate checks.
Common Mistakes
- Do not pair modalities by file order or display name.
- Do not treat derived OCR as exempt from the source asset’s permission.
- Do not allow a target label to enter the input text.
