Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Project: detect conveyor jams from timestamped video clips

Last updated: 7 Oct 20265 min read
project
AdvancedBy AITrove Editorial

Train a short-clip event model and test it by independent recording, first-alert delay and false pages per camera-hour.

Build the event ledger

Record camera and conveyor IDs, physical recording IDs, frame timestamps and reviewed jam start and end times. Extract clips ending at each proposed live decision time; label them by a declared event-overlap rule. Split whole recordings before clip extraction, then reserve later time periods or new camera angles for a final transfer test. Keep ambiguous onset frames available for review. The clip lesson prevents adjacent-clip leakage.

Run the model wiring check

The code performs one synthetic update on a tiny 3D convolutional classifier. It proves batch, channel, frame, height and width ordering plus a finite class loss; it cannot establish real event detection. Train the actual model against a one-frame baseline and a motion-statistic baseline under the same event ledger. Test a frame-order shuffle to see whether the candidate depends on temporal evidence. The temporal lesson explains aggregation tradeoffs.

Score independent events

Merge neighboring positive clip scores into one alert with a frozen cooldown. Report jam-event recall, first-alert delay, false pages per camera-hour and recall by jam duration. Group bootstrap or uncertainty calculations by recording rather than by thousands of overlapping clips. Inspect lighting transitions, maintenance, camera shakes and temporarily stopped belts. A high clip accuracy can coexist with repeated false pages during a single maintenance period.

Test actual live replay

Replay a recording in timestamp order and make each decision using only frames available by that time. Include video decode, resize, clip buffering, model forward and cooldown in p95 latency. Test dropped frames and missing video: output an unavailable status or fallback signal rather than a confident no-jam. Select thresholds on development recordings; do not watch the final replay and retune.

Release with a clear rollback

Package sampling times, resize, normalization, model weights, threshold, cooldown and camera-specific eligibility. Gate on event recall, false-page budget and first-alert deadline. Retain the previous operational monitor if the model misses any gate. The deliverable includes event-level counts, replay traces and failure clips with reviewed timestamps, not just one averaged score.

Implementation

python
import torch
from torch import nn
from torch.nn import functional as functional

torch.manual_seed(47)
timestamped_clips = torch.rand(4, 1, 5, 8, 8)
reviewed_jam_labels = torch.tensor([0, 1, 0, 1])
jam_model = nn.Sequential(nn.Conv3d(1, 8, kernel_size=3, padding=1),
                          nn.ReLU(), nn.AdaptiveAvgPool3d(1),
                          nn.Flatten(), nn.Linear(8, 2))
optimizer = torch.optim.AdamW(jam_model.parameters(), lr=0.0007)
optimizer.zero_grad(set_to_none=True)
jam_logits = jam_model(timestamped_clips)
assert jam_logits.shape == (4, 2)
loss = functional.cross_entropy(jam_logits, reviewed_jam_labels)
loss.backward()
optimizer.step()
assert torch.isfinite(loss)

Performance and operating cost

Clip input and intermediate activations scale with BTHW for B clips, T frames and H-by-W resolution, multiplied by model channel widths. Serving cost includes repeated decode and overlapping clip inference; event cooldown reduces alert count but not necessarily compute. Increasing frame rate or clip overlap can improve onset coverage while raising throughput needs. The toy update is a tensor-path check, not a measurement of event recall or camera-hour false pages.

Common Mistakes

  • Do not score overlapping positive clips as separate independent jams.
  • Do not use future frames in a claimed live decision.
  • Do not hide dropped video behind a negative class prediction.

Read next

ai-data
deep-learning
Storage details