A waveform model receives samples at a declared rate; a spectrogram adds window, hop and padding choices that must match training and serving.
Audio sample rates, short-time spectra and window contracts
Start with physical time
A recording of 32,000 samples at 16,000 samples per second lasts two seconds. The same array tagged as 8,000 samples per second appears four seconds long and shifts every frequency interpretation. Keep original rate, target rate, channel layout and recording-device identity in the manifest. Resampling means calculating new samples with an appropriate anti-aliasing filter; changing metadata is not resampling. The code checks duration and the spectral frame layout for one consistent rate. Sequence lengths later depend on this contract.
Declare the analysis window
A short-time Fourier transform divides a waveform into overlapping windows. The transform’s FFT width controls frequency-bin spacing, while the hop controls how often a frame begins. With centered padding disabled, a 32,000-sample signal, 400-sample window and 160-sample hop yields 198 complete frames under a standard floor-count rule. A centered transform yields a different count because it pads the ends. State the window function, padding mode, return type and magnitude or power conversion; a tensor with the expected rank can still have the wrong time base.
Protect the label interval
A machine alarm label may mark a brief squeal inside a long recording. Training on a clip labeled alarm when the selected crop excludes that squeal creates incorrect supervision. Keep event start and end times, then decide whether a window is positive by overlap, center time or any audible presence. The rule affects class balance and false-alert interpretation. Save waveform excerpts and spectrogram snapshots for reviewed cases. Masking must preserve the evidence needed for the label.
Normalize with training-only statistics
Record microphone gain, clipping fraction, channel mixing and silence handling. A log-power spectrum usually needs a floor before logarithms so silent bins remain finite. Any mean and scale applied to features should come from training recordings, then be frozen for later devices and sites. A model can learn microphone hiss or compression artifacts rather than a fault. Group the split by physical machine and recording session before extracting overlapping windows. The split boundary applies to every crop.
Compare deployed preprocessing byte for byte
Replay a fixed waveform through both training and serving paths and compare sample count, spectral bins, frame count, value range and output logits within tolerance. Check p95 time from audio capture through resampling, feature extraction and model scoring. A fast neural forward does not compensate for a slow decoder or resampler. The applied project gates the full alert path and review load.
Implementation
import torch
sample_rate = 16000
sample_count = 32000
window_size = 400
hop_size = 160
waveform = torch.linspace(-0.7, 0.7, sample_count)
window = torch.hann_window(window_size)
complex_spectrum = torch.stft(waveform, n_fft=window_size,
hop_length=hop_size, win_length=window_size,
window=window, center=False, return_complex=True)
power_spectrum = complex_spectrum.abs().square()
log_power = torch.log(power_spectrum.clamp_min(1e-7))
assert sample_count / sample_rate == 2.0
assert log_power.shape == (window_size // 2 + 1,
1 + (sample_count - window_size) // hop_size)
assert torch.isfinite(log_power).all()Performance and operating cost
For N samples, FFT width F and hop H, the frame count is approximately N/H, and a direct FFT implementation uses roughly O((N/H) F log F) arithmetic with O((N/H) F) spectral storage. A log-power representation adds no trainable parameters but can become a large intermediate tensor for long recordings. Resampling and decode add separate costs. The example verifies frame geometry and finite values, not alarm accuracy or correct resampling of a real file.
Common Mistakes
- Do not relabel sample-rate metadata and call that resampling.
- Do not infer event presence from a recording-level label after cropping.
- Do not change centered padding or hop size between training and serving.
