Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Invertible flows and change-of-variables accounting

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

An invertible model assigns density by mapping observations to a known base distribution and adding the log volume change of every transformation.

Start with continuous, owned features

Suppose a service window has two reviewed continuous features: standardized request latency and standardized error pressure. A flow maps this vector to a base distribution while retaining enough information to invert the mapping. This is different from an autoencoder that may compress dimensions or a classifier that predicts a label directly. Keep feature units, normalization statistics and missing-value policy fixed. A density from one transformed feature scale is not numerically comparable with a density from another scale without accounting for the transformation.

Carry the Jacobian term

If z is the transformed observation, the data log density equals base log density at z plus the log absolute determinant of the forward data-to-base Jacobian. Omitting that determinant rewards arbitrary shrinking or stretching. For a triangular affine coupling, only the scaled coordinates contribute to the determinant, so it can be computed without building a dense matrix. The code transforms a two-feature vector, inverts it and checks a finite log density. The coupling lesson extends the layer.

Separate exact math from model usefulness

An exact likelihood calculation means the formula is tractable for the chosen continuous representation. It does not mean high likelihood is equivalent to operational normality, nor that rare incidents always receive lower likelihood. A flow can spend capacity on irrelevant pixel or telemetry detail. Compare its alert ranking with a simple baseline and stratify by incident type. The reconstruction project provides a different anomaly score for comparison.

Keep inversion testable

Every transform in a chain must have a declared inverse and log-determinant sign. Compose forward log determinants by addition, and test reconstruction error on fixed examples and extreme but valid feature values. A forward transform using z equals x minus shift times an exponential has the opposite sign from the inverse transform; a sign error may still produce finite losses. Check the analytic density on a tiny case against a direct calculation before training a large model.

Plan the serving boundary

Persist base-distribution parameters, feature standardization, layer order and alert threshold with the weights. Score only windows built from measurements observable by the alert time. A missing feature should trigger a documented fallback instead of an arbitrary zero that could look highly probable. The project evaluates alerts by reviewed incident and on-call workload, rather than treating likelihood alone as a release gate.

Implementation

python
from math import exp, isclose, log, pi

latency_signal, error_signal = 2.0, 3.0
log_scale = 0.2 * latency_signal
shift = 0.5 * latency_signal
base_first = latency_signal
base_second = (error_signal - shift) * exp(-log_scale)
forward_log_determinant = -log_scale
base_log_density = -0.5 * (base_first ** 2 + base_second ** 2) - log(2 * pi)
observed_log_density = base_log_density + forward_log_determinant
restored_error_signal = base_second * exp(log_scale) + shift
assert isclose(restored_error_signal, error_signal)
assert isclose(observed_log_density,
               -0.5 * (4.0 + (2.0 * exp(-0.4)) ** 2) - log(2 * pi) - 0.4)

Performance and operating cost

A two-feature affine coupling is O(1). For D features and L layers, forward work depends on each conditioner network plus O(LD) elementwise transforms; saved activations and conditioner widths dominate training memory. A dense Jacobian determinant would be much more expensive, which is why structured coupling matters. Exact density does not remove the cost of fitting, calibrating or reviewing anomalies. The example is a mathematical self-check, not a trained incident model.

Common Mistakes

  • Do not omit or reverse the forward log-determinant sign.
  • Do not compare raw density values across different feature scales.
  • Do not equate low density with a confirmed incident.

Read next

ai-data
deep-learning
Storage details