A coupling layer changes one feature group using the other, allowing a cheap inverse and triangular log determinant; repeated layers need controlled scale and feature mixing.
Affine coupling layers, inverse checks and scale stability
Split features deliberately
An affine coupling leaves one feature group unchanged while a conditioner predicts shift and log scale for the other group. This makes inversion direct even if the conditioner itself is a noninvertible neural network. In a two-feature illustration, latency conditions the transformation of error pressure. If every layer always leaves latency untouched, the model never reshapes that dimension; alternate or permute feature groups across layers. Keep the permutation recorded so an inverse uses the exact reverse order. Density accounting supplies the sign convention.
Clamp the learned log scale
An unconstrained exponential scale can overflow, underflow or produce extreme density terms. Bound the log scale with a smooth function and monitor its distribution. The code uses a bounded hyperbolic tangent; the bound is an engineering choice to validate on real standardized telemetry, not a universal constant. A bounded scale can also limit expressiveness, so inspect held-out likelihood and inverse error. Use finite-value checks before an optimizer step to avoid corrupting model state.
Check both directions
For forward data-to-base mapping, subtract shift and multiply by the negative exponential scale; its log determinant is negative scale. The inverse multiplies by the positive exponential and adds shift. Test round-trip reconstruction per batch and compare forward plus inverse log determinants with zero. A mistaken feature permutation can preserve shapes and even train while making inversion wrong. The miniature code asserts exact tensor closeness for one layer.
Handle missing and discrete inputs
A continuous density model does not automatically assign a probability mass to discrete counts. Counts may need a separate model or a documented continuous representation; naive jitter changes the target distribution. Missing values are not ordinary zeros after standardization. Segment by completeness and route incomplete windows through a separately evaluated path. A feature mask added to the vector may be useful, but then train and serve with that mask and audit missingness drift.
Validate against operations
Likelihood should be assessed on unseen time periods and service families, not only by average negative log likelihood. Review false alerts during deployments, quiet nights and high-traffic but healthy periods. Compare with an autoencoder and simple percentile rule at equal alert budgets. The applied project freezes an on-call threshold only after the model passes inverse, density and workload checks.
Implementation
import torch
from torch import nn
torch.manual_seed(47)
telemetry = torch.tensor([[0.8, -0.3], [1.4, 0.2], [-0.5, 0.9]])
conditioner = nn.Linear(1, 2)
shift_and_raw_scale = conditioner(telemetry[:, :1])
shift = shift_and_raw_scale[:, :1]
log_scale = shift_and_raw_scale[:, 1:].tanh() * 0.8
base_second = (telemetry[:, 1:] - shift) * torch.exp(-log_scale)
base_vector = torch.cat((telemetry[:, :1], base_second), dim=1)
forward_log_determinant = -log_scale.squeeze(1)
restored_second = base_vector[:, 1:] * torch.exp(log_scale) + shift
restored = torch.cat((base_vector[:, :1], restored_second), dim=1)
torch.testing.assert_close(restored, telemetry)
torch.testing.assert_close(forward_log_determinant + log_scale.squeeze(1),
torch.zeros(3))Performance and operating cost
The affine elementwise transform is O(BD) for batch B and transformed dimension D, plus the conditioner network cost. The triangular determinant requires summing log scales, not materializing a D-by-D Jacobian. Multiple alternating layers add conditioner work and backward activations. The small example verifies one layer’s inverse and determinant signs; it does not establish whether the resulting density is useful for incident ranking. Profile scale extremes and inverse error on the actual feature distribution.
Common Mistakes
- Do not leave the same feature group unchanged in every layer.
- Do not let an exponential log scale grow without finite-value and inverse checks.
- Do not treat missing telemetry as a normal standardized zero without a policy.
