Combining image and text signals should beat strong single-channel baselines while remaining defined when one channel fails.
Multimodal fusion: choose a baseline and handle absent channels
Start with separate models
Measure an image-only and a text-only receipt classifier before joining their signals. A late-fusion baseline can average calibrated scores with fixed weights, which makes each contribution inspectable. Early feature fusion may capture interactions, but it also couples preprocessing and can amplify a mislabeled pair. Pair quality limits every fusion method.
Define missing behavior
OCR may fail on a folded receipt; an image may be withheld while approved text remains. Specify image-only, text-only and no-evidence routes explicitly. Do not fill a missing embedding with zeros and call it a genuine observation unless the model was trained and tested with that convention. Record modality availability and fallback rate by document type.
Check calibration and dominance
If one model produces scores near 0.99 and another near 0.6, a naive average can let scale rather than evidence determine the result. Calibrate on held-out data, inspect disagreement cases and compare fusion against the stronger single-modality model. Calibration checks matter before an alert threshold uses the combined score.
Rehearse failure modes
Use a receipt with a clear image but empty OCR, another with legible text but a blocked image, and a third whose image and text refer to different receipt IDs. The first two should take declared fallbacks; the mismatched pair should be rejected before scoring. Report cost and latency for every route, since running two encoders can exceed the request budget.
Implementation
def late_fusion(image_score, text_score, image_weight=0.55):
if image_score is None:
return text_score
if text_score is None:
return image_score
return image_weight * image_score + (1 - image_weight) * text_scorePerformance and operating cost
Score fusion is O(1) time and space after encoders run. Image and text encoding often dominate latency and memory; a missing-channel fallback can save work, but only if routing happens before the unused encoder is invoked.
Common Mistakes
- Do not assume a zero vector means the same thing as missing media.
- Do not celebrate fusion without comparing it with each single-channel baseline.
- Do not average uncalibrated scores and infer equal evidence.
