An autoencoder learns to reconstruct accepted normal windows; its reconstruction error becomes an anomaly score only after missing values and threshold calibration are defined.
Masked reconstruction loss and anomaly thresholds
Define what the model sees
An encoder compresses a telemetry window into a smaller representation and a decoder predicts the original features. Training on a pool labeled normal does not guarantee every row is normal. Audit contamination from known faults, maintenance events and sensor outages before fitting. Normalize each feature with statistics computed only from the training population, and persist those statistics for serving. Split discipline keeps threshold calibration separate from fitting.
Mask missing measurements
Zero may be a meaningful reading, so do not use zero as a missing-value marker. Store a Boolean observation mask and impute absent inputs from training-set statistics before the encoder. Calculate reconstruction loss only on observed targets, dividing by the observed count per window. A window with no observed fields has no defensible score and should be rejected. Count coverage beside score, since a low error based on one field is not comparable with a full six-field window.
Choose the scoring unit
For each window, squared error can be averaged over observed feature positions. That mean gives a comparable scale across different coverage, but it may hide one critical sensor spike under five quiet sensors. Consider separate feature-family scores or a maximum standardized residual when the product needs sensitivity to localized faults. Pin the choice before threshold tuning. A decoder with too much capacity may reconstruct faults well, so low error is not a proof of normality. The alert project tests this failure mode.
Calibrate without reusing training loss
Fit the model on training normals, score a disjoint normal calibration set and choose a threshold under an explicit alert-budget target. A percentile on four reference windows is only a code illustration, not a credible operating threshold. Use enough independent device-days to estimate the tail, and report the numerator and denominator for false alarms. If labeled incidents exist, inspect detection recall on a separate incident set, but do not tune repeatedly on the final test group.
Monitor shifts after release
A sensor replacement or seasonal operating mode can raise reconstruction error without a fault. Track feature coverage, score distribution, alert rate and downstream confirmed incidents by device cohort. Store model, preprocessing and threshold revisions together so a rollback restores the decision rule, not just weights. A changed score distribution calls for investigation and possibly fresh validation, not an automatic threshold increase that hides genuine failures.
Implementation
import torch
from torch import nn
class TelemetryAutoencoder(nn.Module):
def __init__(self):
super().__init__()
self.encoder = nn.Sequential(nn.Linear(6, 4), nn.ReLU(), nn.Linear(4, 2))
self.decoder = nn.Sequential(nn.Linear(2, 4), nn.ReLU(), nn.Linear(4, 6))
def forward(self, readings: torch.Tensor) -> torch.Tensor:
return self.decoder(self.encoder(readings))
torch.manual_seed(73)
normal_readings = torch.tensor([[0.2, 0.4, 0.1, 0.7, 0.3, 0.6],
[0.3, 0.5, 0.2, 0.6, 0.4, 0.0]])
observed = torch.tensor([[True, True, True, True, True, True],
[True, True, True, True, True, False]])
training_medians = torch.tensor([0.25, 0.45, 0.15, 0.65, 0.35, 0.55])
model_inputs = torch.where(observed, normal_readings, training_medians)
reconstruction = TelemetryAutoencoder()(model_inputs)
observed_count = observed.sum(dim=1)
if bool((observed_count == 0).any()):
raise ValueError("unscorable telemetry window")
window_scores = (((reconstruction - normal_readings).square() * observed)
.sum(dim=1) / observed_count)
assert window_scores.shape == (2,)
assert torch.isfinite(window_scores).all()Performance and operating cost
For input width D and hidden widths H and Z, dense encoder/decoder work scales with the matrix products across their layers; parameter storage does not grow with the number of windows. Per-window masked score calculation is O(D). Calibration requires scoring a separate population and retaining enough score samples or a controlled quantile sketch to estimate the tail. A larger decoder can improve reconstruction while making anomalies less separable.
Common Mistakes
- Do not turn missing readings into genuine zeros in the target loss.
- Do not choose an alert threshold from the same training windows used to fit the model.
- Do not assume every high-error window is a fault or every low-error window is safe.
