Train a short-clip event model and test it by independent recording, first-alert delay and false pages per camera-hour.
Project: detect conveyor jams from timestamped video clips
Build the event ledger
Record camera and conveyor IDs, physical recording IDs, frame timestamps and reviewed jam start and end times. Extract clips ending at each proposed live decision time; label them by a declared event-overlap rule. Split whole recordings before clip extraction, then reserve later time periods or new camera angles for a final transfer test. Keep ambiguous onset frames available for review. The clip lesson prevents adjacent-clip leakage.
Run the model wiring check
The code performs one synthetic update on a tiny 3D convolutional classifier. It proves batch, channel, frame, height and width ordering plus a finite class loss; it cannot establish real event detection. Train the actual model against a one-frame baseline and a motion-statistic baseline under the same event ledger. Test a frame-order shuffle to see whether the candidate depends on temporal evidence. The temporal lesson explains aggregation tradeoffs.
Score independent events
Merge neighboring positive clip scores into one alert with a frozen cooldown. Report jam-event recall, first-alert delay, false pages per camera-hour and recall by jam duration. Group bootstrap or uncertainty calculations by recording rather than by thousands of overlapping clips. Inspect lighting transitions, maintenance, camera shakes and temporarily stopped belts. A high clip accuracy can coexist with repeated false pages during a single maintenance period.
Test actual live replay
Replay a recording in timestamp order and make each decision using only frames available by that time. Include video decode, resize, clip buffering, model forward and cooldown in p95 latency. Test dropped frames and missing video: output an unavailable status or fallback signal rather than a confident no-jam. Select thresholds on development recordings; do not watch the final replay and retune.
Release with a clear rollback
Package sampling times, resize, normalization, model weights, threshold, cooldown and camera-specific eligibility. Gate on event recall, false-page budget and first-alert deadline. Retain the previous operational monitor if the model misses any gate. The deliverable includes event-level counts, replay traces and failure clips with reviewed timestamps, not just one averaged score.
Implementation
import torch
from torch import nn
from torch.nn import functional as functional
torch.manual_seed(47)
timestamped_clips = torch.rand(4, 1, 5, 8, 8)
reviewed_jam_labels = torch.tensor([0, 1, 0, 1])
jam_model = nn.Sequential(nn.Conv3d(1, 8, kernel_size=3, padding=1),
nn.ReLU(), nn.AdaptiveAvgPool3d(1),
nn.Flatten(), nn.Linear(8, 2))
optimizer = torch.optim.AdamW(jam_model.parameters(), lr=0.0007)
optimizer.zero_grad(set_to_none=True)
jam_logits = jam_model(timestamped_clips)
assert jam_logits.shape == (4, 2)
loss = functional.cross_entropy(jam_logits, reviewed_jam_labels)
loss.backward()
optimizer.step()
assert torch.isfinite(loss)Performance and operating cost
Clip input and intermediate activations scale with BTHW for B clips, T frames and H-by-W resolution, multiplied by model channel widths. Serving cost includes repeated decode and overlapping clip inference; event cooldown reduces alert count but not necessarily compute. Increasing frame rate or clip overlap can improve onset coverage while raising throughput needs. The toy update is a tensor-path check, not a measurement of event recall or camera-hour false pages.
Common Mistakes
- Do not score overlapping positive clips as separate independent jams.
- Do not use future frames in a claimed live decision.
- Do not hide dropped video behind a negative class prediction.
