Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Video clip timestamps and recording-group evaluation

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Short-video labels depend on frame timestamps and source-recording identity; splitting adjacent clips independently can make held-out accuracy meaningless.

Sample by elapsed time

A camera may drop frames or vary its frame rate. Taking every fifth decoded frame then changes the real-time span across recordings. Choose target timestamps relative to the clip start and select or interpolate frames according to a declared policy. The code selects frames nearest 0, 120 and 240 milliseconds from an irregular timestamp list. A duplicate selected frame or a large timestamp gap should be recorded rather than silently treated as smooth motion. The model lesson expects a fixed frame count.

Attach labels to intervals

A conveyor jam label may describe seconds four through six of a ten-second recording. A clip from seconds zero through two is not a jam clip just because it came from the same file. Define overlap or center-time rules for assigning labels and mark ambiguous boundary clips for review. Keep source recording ID, camera, line, shift and event interval with every derived clip. Save timestamped thumbnails to audit whether selected frames actually contain the action.

Split at the source boundary

Many overlapping clips from one recording share most pixels and background. A random clip split lets near-duplicates into train and test, and a model can memorize the conveyor or camera angle. Split by recording, then consider a stricter camera or site holdout for transfer claims. Report the number of independent recordings and jams, not merely the inflated clip count. Augmentation leakage has the same parent-identity problem.

Preserve temporal order and sampling revision

A model with a temporal convolution receives frames in a specific order. Reversing them can change the meaning of a jam onset even if average appearance is similar. Record target times, clip duration, frame resize and any missing-frame substitution. If inference uses a rolling stream while training used offline clips centered around an event, the model may have seen future frames during training. Declare whether the task is retrospective tagging or live early warning.

Evaluate at event level

Adjacent positive clips from one jam should count as one incident for recall and alert load. Report event recall, first-alert delay, false alerts per camera-hour, and performance by lighting and camera movement. Compare against a static-frame baseline to prove the temporal model contributes more than background recognition. The project ties those metrics to a reversible release.

Implementation

python
def nearest_timestamp_indices(frame_times_ms, target_times_ms):
    if not frame_times_ms or frame_times_ms != sorted(frame_times_ms):
        raise ValueError("frame timestamps must be ordered")
    return [min(range(len(frame_times_ms)),
                key=lambda index: (abs(frame_times_ms[index] - target), index))
            for target in target_times_ms]

decoded_frame_times = [0, 40, 83, 121, 160, 203, 241]
selected = nearest_timestamp_indices(decoded_frame_times, [0, 120, 240])
assert selected == [0, 3, 6]
assert [decoded_frame_times[index] for index in selected] == [0, 121, 241]

Performance and operating cost

This direct search is O(FT) time for F decoded frames and T requested timestamps, with O(T) output space; a two-pointer pass over sorted timestamps can reduce selection to O(F + T) for long recordings. Decoding and storing all frames usually cost far more memory and time. A clip with C channels, T frames and H-by-W resolution uses O(CTHW) input values before model activations. Record actual time gaps so a fixed tensor shape does not conceal variable motion speed.

Common Mistakes

  • Do not split adjacent clips from one recording between training and final test.
  • Do not assign a full-recording event label to clips outside the event interval.
  • Do not treat frame index spacing as fixed elapsed time on variable-rate video.

Read next

ai-data
deep-learning
Storage details