Build an escalation classifier for variable-length histories, then prove that padding, customer mixing and carried state cannot create false evidence.
Project: predict service escalation from event histories
Freeze the event contract
Represent each service event by a timestamp, event type and known-at-event features. Exclude notes or outcomes recorded after the prediction time. Split 470 histories by customer and ticket before windowing; two windows from the same ticket cannot cross train and validation. Keep a label map for routine versus escalated and a count of valid events for each history. Temporal target alignment prevents a model from reading the answer it is supposed to predict.
Start with packed independent histories
Pad a batch only for storage and pack it with measured lengths before an LSTM. Score the last real hidden state with a two-class head. Compare against a simple non-recurrent baseline using the same split and threshold-selection rule. Evaluate by history length and customer cohort so an apparent gain on long tickets cannot hide failures on one-event tickets. The packing contract makes the last event unambiguous.
Add a stateful serving variant carefully
If events arrive one at a time, a state cache can avoid replaying the full history. Key it by ticket identity, version it with the model and set an expiry. Replay and cached-state predictions should agree under deterministic evaluation for a controlled ticket. On handoff to a different ticket, reset both hidden and cell state. A negative test that reuses the previous customer state should fail an isolation assertion. State ownership is a release gate.
Probe leakage and padding
Make a copy of a validation batch and replace only padded rows with extreme values; predictions should remain the same. Swap lengths between two records and confirm the test notices. Then move a post-escalation event into the input to demonstrate how a leaky pipeline can score implausibly well. Report the difference without publishing that leaky model. Keep the counterfactual tests with the data pipeline because a correct network cannot repair a future-looking input.
Release with measured limits
Report customer-grouped precision and recall, false alarms per hundred tickets, latency by history length, peak state-cache occupancy and restart behavior. Save feature encoding, label map, maximum history length, model weights and threshold together. If a ticket exceeds the limit, define whether to replay its latest window or to reject the request. The project is complete when a fresh process can reload the artifact and reproduce the same two fixed-ticket logits.
Implementation
import torch
from torch import nn
from torch.nn.utils.rnn import pack_padded_sequence
class EscalationHistoryModel(nn.Module):
def __init__(self, event_width: int = 4, hidden_width: int = 9):
super().__init__()
self.encoder = nn.LSTM(event_width, hidden_width, batch_first=True)
self.classifier = nn.Linear(hidden_width, 2)
def forward(self, event_features: torch.Tensor,
valid_lengths: torch.Tensor) -> torch.Tensor:
if bool((valid_lengths < 1).any()) or bool((valid_lengths > event_features.size(1)).any()):
raise ValueError("invalid event count")
packed = pack_padded_sequence(event_features, valid_lengths.cpu(),
batch_first=True, enforce_sorted=False)
_, (last_hidden, _) = self.encoder(packed)
return self.classifier(last_hidden[-1])
torch.manual_seed(59)
escalation_model = EscalationHistoryModel().eval()
event_features = torch.randn(3, 5, 4)
valid_lengths = torch.tensor([5, 2, 4])
changed_padding = event_features.clone()
changed_padding[1, 2:] = 1000
changed_padding[2, 4:] = -1000
with torch.inference_mode():
baseline_logits = escalation_model(event_features, valid_lengths)
changed_logits = escalation_model(changed_padding, valid_lengths)
torch.testing.assert_close(baseline_logits, changed_logits)
assert baseline_logits.shape == (3, 2)Performance and operating cost
The LSTM cost grows with valid event count and hidden width, while packing saves work on padding. Its timesteps remain dependent, so a long ticket can dominate latency even in a mixed batch. A stateful service exchanges replay cost for O(active tickets × layers × hidden width) cache memory and stricter identity handling. Measure tail latency and cache eviction, not just average batch throughput.
Common Mistakes
- Do not split windows from one ticket across train and validation.
- Do not let post-escalation events into the input history.
- Do not reuse cached state for a new ticket or silently ignore the supplied valid lengths.
Read next
- Packed LSTM sequences and the last valid state
- Recurrent hidden-state resets and truncated gradients
- Causal target shifts and padding-aware loss
- Training and validation modes: measure the model you will serve
- Inference contracts: preserve preprocessing and measure tail latency
Continue the workflow: Project: release an incremental service-event decoder.
Continue the workflow: Project: attribute service incidents with a time-safe graph.
Continue the workflow: Project: forecast warehouse queues with a causal temporal model.
