Detect reviewed fault sounds while preserving sample-rate, event-window and microphone boundaries through training and deployment.
Project: classify machine alarms from short audio windows
Build a recording ledger
Collect machine audio with recording session, site, microphone, rate, channel count and reviewed fault intervals. Split by machine and session before creating overlapping windows. Keep a clean holdout with microphones and operating regimes that were not used to tune thresholds. Define whether the output is window-level fault presence or an event timestamp; the small model below predicts one class per spectral window. The audio lesson fixes the preprocessing geometry.
Run a spectral training check
Extract log-power or log-mel features with frozen settings, normalize using training-only statistics and reject nonfinite or clipped inputs. The code trains a tiny convolutional classifier on already prepared spectral tensors to verify shape, label and gradient flow. It does not read recordings, resample or establish fault performance. Use the exact deployed extraction path for real experiments. Save one fixed waveform and its derived feature tensor as a parity fixture.
Protect short alarms during augmentation
Time or frequency masks must not remove the reviewed fault interval. Separate clean and corrupted evaluation sets, then report detection recall for short and long faults, noise conditions and device families. The masking lesson gives an explicit guard. If a fault event spans several windows, score it as one event after deduplicating neighboring alerts; window accuracy alone can exaggerate success.
Measure on-call cost
Select score and event-merging thresholds on development sessions, then freeze them. Report event recall, median and p95 detection delay, false alarms per machine-hour and review minutes per shift. Compare with a simple energy or spectral rule under the same alarm budget. Inspect false alarms caused by maintenance sounds and benign speed changes. A high window-level score may still page an operator repeatedly for one sustained noise.
Release with missing-audio behavior
Package sample rate, resampling, spectral settings, normalization, window stride, class map, threshold and deduplication policy. Gate on rare-fault recall, false-alarm budget and capture-to-alert latency. If a microphone disconnects or produces invalid audio, report unavailable input instead of a confident healthy prediction. Replay fixed recordings in a clean process and retain the earlier alarm path when the candidate misses a gate.
Implementation
import torch
from torch import nn
from torch.nn import functional as functional
torch.manual_seed(47)
log_spectral_windows = torch.rand(4, 1, 16, 20)
reviewed_alarm_classes = torch.tensor([0, 1, 2, 1])
alarm_model = nn.Sequential(nn.Conv2d(1, 8, 3, padding=1), nn.ReLU(),
nn.AdaptiveAvgPool2d(1), nn.Flatten(), nn.Linear(8, 3))
optimizer = torch.optim.AdamW(alarm_model.parameters(), lr=0.0007)
optimizer.zero_grad(set_to_none=True)
alarm_logits = alarm_model(log_spectral_windows)
assert alarm_logits.shape == (4, 3)
loss = functional.cross_entropy(alarm_logits, reviewed_alarm_classes)
loss.backward()
optimizer.step()
assert torch.isfinite(loss)Performance and operating cost
The example consumes O(BFT) spectral values for B windows, F frequency bins and T frames, with convolutional activation storage during training. Real cost includes audio decode, resampling, spectral extraction, overlapping window scoring and alert deduplication. Smaller window strides increase the number of model calls and may reduce detection delay, but can also multiply correlated false alerts. Measure the complete capture-to-alert pipeline on target hardware; the toy update is only a wiring check.
Common Mistakes
- Do not split overlapping windows from one recording across train and test.
- Do not count repeated windows from one fault as independent successful events.
- Do not emit healthy status when the microphone input is unavailable.
