Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Project: attribute service incidents with a time-safe graph

Last updated: 7 Oct 20265 min read
project
AdvancedBy AITrove Editorial

Use a two-hop call graph to rank affected services while auditing edge direction, future-edge leakage, sampled-neighborhood variance and on-call review cost.

Freeze the operational question

For each incident at its first alert, rank services for engineer review. The label is a reviewed affected-service set, not an automatically inferred causal root. Snapshot telemetry and directed dependencies as they were observable at alert time. Keep incident IDs, service IDs, review timestamps and a delayed-label flag. A service with no recorded calls still needs a local-feature score. The message-passing lesson defines the direction convention.

Build and test a minimal graph layer

Represent each service with a fixed-window error-rate and latency feature; normalize using training incidents only. Propagate caller features toward callees for one experiment and reverse edges for another. The code checks a single tensor aggregation step and updates a classifier on three services. It proves that dimensions, aggregation and loss connect; it does not establish incident ranking quality. Require a hand-checked miniature graph before adding relation types or more layers.

Separate seeds from context

Train on incident-service labels for seed nodes only. If the full graph is too large, sample a two-hop context with a fixed fanout and record the draw. Measure whether rare critical upstream dependencies are omitted. Keep the same held-out incidents and feature cutoff for the graph model and the local baseline. The sampling lesson explains why a random node split can overstate transfer.

Evaluate the review workflow

Report recall among reviewed affected services in the first three suggestions, false leads per incident, and time from alert to first correct suggestion. Break results out by new service, isolated service, high-degree shared service and incident size. Inspect cases where reversing edges improves a subset: that may indicate direction-specific failure modes, not permission to choose direction after viewing final test. Include calibration or abstention so an engineer sees when the graph is incomplete. Selective risk is one way to set review coverage.

Ship with graph-version parity

Package the model with graph snapshot revision, feature windows, missing-neighbor rule and fanout. Replay a fixed incident in a clean process and compare scores. Gate release on held-out incident recall, false leads and p95 graph fetch plus scoring latency. Fall back to the local-feature ranker when dependency data is stale. Monitor drift in edge count, hub degree and unavailable services; retraining alone will not repair a broken observability clock.

Implementation

python
import torch
from torch import nn
from torch.nn import functional as functional

torch.manual_seed(47)
service_features = torch.tensor([[0.8, 0.3], [0.4, 0.7], [0.2, 0.5]])
# checkout -> billing, checkout -> inventory; self state is retained.
caller_index = torch.tensor([0, 0, 0, 1, 2])
callee_index = torch.tensor([0, 1, 2, 1, 2])
incoming_sum = torch.zeros_like(service_features)
incoming_sum.index_add_(0, callee_index, service_features[caller_index])
incoming_count = torch.zeros(3, 1)
incoming_count.index_add_(0, callee_index, torch.ones(5, 1))
service_context = incoming_sum / incoming_count
assert service_context.shape == (3, 2)
assert torch.allclose(service_context[1],
                      (service_features[0] + service_features[1]) / 2)

incident_head = nn.Linear(2, 1)
optimizer = torch.optim.AdamW(incident_head.parameters(), lr=0.003)
reviewed_affected = torch.tensor([[0.0], [1.0], [0.0]])
optimizer.zero_grad(set_to_none=True)
loss = functional.binary_cross_entropy_with_logits(
    incident_head(service_context), reviewed_affected)
loss.backward()
optimizer.step()
assert torch.isfinite(loss)

Performance and operating cost

The example materializes E edge messages and aggregates them in O(EF + VF) time for V services and F feature dimensions, with O(VF + EF) feature and message storage before optimizer state. Real incident graphs also pay for timestamped edge retrieval and often for two-hop sampling. Rank evaluation must group service labels by incident; treating every service row as an independent test example inflates certainty. The three-service code is a wiring test, not a benchmark or a claim that caller aggregation is the correct incident direction.

Common Mistakes

  • Do not train on an edge graph assembled after the alert cutoff.
  • Do not score only easy connected services while dropping isolated ones.
  • Do not equate an affected-service label with proven root cause.

Read next

ai-data
deep-learning
Storage details