Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Multimodal fusion: choose a baseline and handle absent channels

Last updated: 7 Oct 20265 min read
tutorial
IntermediateBy AITrove Editorial

Combining image and text signals should beat strong single-channel baselines while remaining defined when one channel fails.

Start with separate models

Measure an image-only and a text-only receipt classifier before joining their signals. A late-fusion baseline can average calibrated scores with fixed weights, which makes each contribution inspectable. Early feature fusion may capture interactions, but it also couples preprocessing and can amplify a mislabeled pair. Pair quality limits every fusion method.

Define missing behavior

OCR may fail on a folded receipt; an image may be withheld while approved text remains. Specify image-only, text-only and no-evidence routes explicitly. Do not fill a missing embedding with zeros and call it a genuine observation unless the model was trained and tested with that convention. Record modality availability and fallback rate by document type.

Check calibration and dominance

If one model produces scores near 0.99 and another near 0.6, a naive average can let scale rather than evidence determine the result. Calibrate on held-out data, inspect disagreement cases and compare fusion against the stronger single-modality model. Calibration checks matter before an alert threshold uses the combined score.

Rehearse failure modes

Use a receipt with a clear image but empty OCR, another with legible text but a blocked image, and a third whose image and text refer to different receipt IDs. The first two should take declared fallbacks; the mismatched pair should be rejected before scoring. Report cost and latency for every route, since running two encoders can exceed the request budget.

Implementation

python
def late_fusion(image_score, text_score, image_weight=0.55):
    if image_score is None:
        return text_score
    if text_score is None:
        return image_score
    return image_weight * image_score + (1 - image_weight) * text_score

Performance and operating cost

Score fusion is O(1) time and space after encoders run. Image and text encoding often dominate latency and memory; a missing-channel fallback can save work, but only if routing happens before the unused encoder is invoked.

Common Mistakes

  • Do not assume a zero vector means the same thing as missing media.
  • Do not celebrate fusion without comparing it with each single-channel baseline.
  • Do not average uncalibrated scores and infer equal evidence.

Read next

ai-data
multimodal-ai
Storage details