Connect image and text evidence with validated pairs, compatible representations, missing-channel fallbacks and traceable evaluation.
Learning path
- Multimodal data contract: paired records, identity and consent
- Media preprocessing: image geometry, text spans and time alignment
- Cross-modal representation: positive pairs, negatives and shared space
- Multimodal fusion: choose a baseline and handle absent channels
- Cross-modal retrieval: candidate pools, ranking and versioned indexes
- Multimodal evaluation: subgroup failures and evidence traceability
- Project: audit an image-and-text receipt matcher
Connected foundations
Use the earlier data and model boundaries as prerequisites.
