Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Multimodal data contract: paired records, identity and consent

Last updated: 5 Oct 20265 min read
tutorial
IntermediateBy AITrove Editorial

A multimodal example is a governed relationship among media items, labels and a target decision, not merely files sharing a directory.

Define the pair

AI Trove could link a receipt image to its extracted text and a reviewer’s description. Give each asset an immutable media ID, parent receipt ID, capture time, source and transform version. A filename is not an identity contract: resizes and retries can produce several files for one receipt. Image-unit rules should determine whether the training unit is a receipt, page or crop.

Record alignment

The text must correspond to the same receipt image and, for multipage documents, the same page or region. Store alignment method and confidence; do not silently pair records by row order after a missing upload. A corrected OCR result should keep a link to the earlier version so a historical prediction can be reconstructed. The target label must stay distinct from the text used as model input.

Respect availability and permission

A user may permit a receipt image for a specific workflow but not indefinite training reuse. Track permitted purpose, retention deadline and deletion state for each asset and derived representation. A text transcript can contain the same sensitive fields as the image. Inference logging must obey those limits, including when a model emits intermediate descriptions.

Test a broken pair

Create image I-47 with OCR T-47, then insert a late text row T-48 for a different receipt. A positional join should fail the fixture; a keyed join should retain the correct pair. Add a deleted image whose cached text remains, and require the resulting training example to be excluded until the deletion policy is resolved.

Implementation

python
def paired_receipts(images, texts):
    text_by_receipt = {row["receipt_id"]: row for row in texts}
    return [(image, text_by_receipt[image["receipt_id"]])
            for image in images
            if image["receipt_id"] in text_by_receipt
            and image["consent_ok"] and text_by_receipt[image["receipt_id"]]["consent_ok"]]

Performance and operating cost

A keyed join costs O(I + T) expected time and O(T + P) memory for I images, T text rows and P accepted pairs. The simple map assumes one text version per receipt; production joins need version and page keys plus duplicate checks.

Common Mistakes

  • Do not pair modalities by file order or display name.
  • Do not treat derived OCR as exempt from the source asset’s permission.
  • Do not allow a target label to enter the input text.

Read next

ai-data
multimodal-ai
Storage details