Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Project: review telemetry anomalies with an invertible flow

Last updated: 7 Oct 20265 min read
project
AdvancedBy AITrove Editorial

Fit a continuous density model on healthy service windows and test whether its low-density alerts improve reviewed incident recall at a fixed on-call budget.

Define a healthy training population

Use standardized continuous latency and error-pressure features from reviewed healthy windows. Split by time and service family, and isolate deployments or incidents from the healthy fit. Record feature arrival times so scoring never sees later telemetry. Store normalization statistics from training only. The synthetic code below trains one two-feature coupling layer and checks finite density; it does not prove anomaly-detection performance. The density lesson defines the calculation.

Train with inverse tests attached

Alternate feature order across enough layers to transform all dimensions. After every architecture change, run forward–inverse round trips, determinant-sign checks and finite-loss tests at normal and extreme valid inputs. Inspect per-feature standardization and learned log-scale range. A lower negative log likelihood on healthy windows may merely improve modeling of routine traffic variation. The coupling lesson explains why the inverse and scale must be audited.

Calibrate an alert policy

Turn low density into an alert only after selecting a threshold on a development period with a fixed daily review budget. Compare a simple percentile rule and the existing reconstruction model under the same alert budget. Evaluate reviewed incident recall, false pages per service-day and time-to-first-alert by incident type. Keep uncertainty about unlabeled incidents visible. A surprising but healthy deployment can be low-density, while a harmful common-looking incident can be high-density.

Test shifts and missing data

Score unseen service families and later periods. Break results down by high traffic, maintenance, telemetry gaps and feature latency. Route missing or nonfinite inputs to a stated fallback rather than passing them into the flow. Preserve alert examples with feature values, base coordinates and density components for review, while avoiding claims that a low-density feature caused an incident. A human incident label and a model anomaly score answer different questions.

Release with rollback and workload gates

Package feature units, normalization, layer order, base distribution, model weights and threshold. Replay fixed healthy and incident windows in a clean process. Gate on incident recall at a maximum false-page rate, finite density, inverse error and p95 scoring latency. Keep the simpler monitor when the flow misses a gate. The deliverable is an alert ledger tied to review capacity, not a histogram of impressive likelihood values.

Implementation

python
import math
import torch
from torch import nn

torch.manual_seed(47)
healthy_windows = torch.tensor([[0.2, -0.3], [0.5, 0.1],
                                [-0.4, 0.2], [0.1, -0.2]])
conditioner = nn.Linear(1, 2)
optimizer = torch.optim.AdamW(conditioner.parameters(), lr=0.0007)
optimizer.zero_grad(set_to_none=True)
shift_and_scale = conditioner(healthy_windows[:, :1])
shift = shift_and_scale[:, :1]
log_scale = shift_and_scale[:, 1:].tanh() * 0.8
latent_second = (healthy_windows[:, 1:] - shift) * torch.exp(-log_scale)
latent = torch.cat((healthy_windows[:, :1], latent_second), dim=1)
base_log_density = -0.5 * latent.square().sum(dim=1) - math.log(2 * math.pi)
data_log_density = base_log_density - log_scale.squeeze(1)
loss = -data_log_density.mean()
assert torch.isfinite(loss)
loss.backward()
optimizer.step()

Performance and operating cost

One two-feature coupling layer is cheap, but it leaves one coordinate unchanged and is only a training-path demonstration. A usable flow stacks alternating conditioners, whose computation and saved activations dominate O(BD) elementwise transforms. Scoring every service window also incurs feature-fetch and logging costs. Review labels, incident grouping and the agreed alert budget determine whether a gain matters. The synthetic optimizer step verifies a finite gradient, not a validated anomaly detector.

Common Mistakes

  • Do not claim that a finite synthetic likelihood proves useful alerts.
  • Do not fit normalization on future or incident windows.
  • Do not deploy low-density pages without a false-page budget and missing-data fallback.

Read next

ai-data
deep-learning
Storage details