Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Media preprocessing: image geometry, text spans and time alignment

Last updated: 5 Oct 20265 min read
tutorial
IntermediateBy AITrove Editorial

Preprocessing must preserve the coordinates and timing needed to explain which evidence came from which media item.

Keep coordinate transforms

If a receipt image is rotated, cropped or resized, save the transform so OCR boxes can be mapped back to original pixels. A model may read the correct characters yet attach them to the wrong line after a coordinate mistake. Pixel geometry should be tested with corner points and boxes, not only a visual preview.

Normalize text deliberately

Record OCR engine version, Unicode normalization and tokenization policy. Do not strip currency symbols, decimal separators or negative signs from payment totals. Preserve raw and normalized text separately, with offsets that can point to the source image. A normalization change can shift token boundaries and invalidate region-level annotations even when the visible sentence appears similar.

Align asynchronous streams

For audio or video, timestamps connect words, frames and speaker turns. Specify clock origin and allowed skew; one second of drift can move a spoken instruction onto the wrong frame. Missing segments should remain explicit. A blank transcript is different from a verified silent segment, and a duplicated frame must not become independent evidence.

Run a fixture

Use a receipt image rotated 90 degrees with a total box near the bottom edge, plus OCR text whose amount is 47.80. Map the transformed box back and verify it overlaps the original amount. Add a missing text segment and one OCR correction; the training row should identify which text version was available at prediction time.

Implementation

python
def shift_media_times(media_rows, clock_offset_ms):
    return [{**row, "aligned_ms": row["source_ms"] + clock_offset_ms}
            for row in media_rows]

Performance and operating cost

Timestamp alignment is O(N) time and O(N) output space for N segments. Image warps cost proportional to pixel count; storing both original and transformed coordinates increases metadata but enables audit and error localization.

Common Mistakes

  • Do not discard geometry after resizing or rotating an image.
  • Do not normalize away characters that determine a financial amount.
  • Do not interpret missing text as confirmed silence.

Read next

ai-data
multimodal-ai
Storage details