Train an encoder-decoder model on time-bounded incident records, then gate release on source support, severe-fact recall and full-decoding behavior.
Project: generate grounded summaries of service incidents
Build a source-and-summary ledger
Each record needs incident ID, source event IDs, event timestamps, decision cutoff and a reviewed summary with links to supporting events. Remove notes written after the cutoff. Split by whole incident and reserve a later period or new service team. Store source and target tokenizer revisions. The cross-attention lesson fixes source scope and padding.
Fit against an extractive baseline
First make a deterministic summary from event fields and compare it with the proposed neural model. Train with shifted targets, separate source and target masks, and loss that ignores pad labels. The code exercises one synthetic encoder-decoder forward and loss; its random token IDs do not establish factual generation. The shift lesson explains why the target input omits the token being predicted.
Review actual generations
Decode complete summaries on held-out incidents. Mark unsupported statements, severe-event omissions, repeated claims and premature stopping. Report these rates by incident type and source length, alongside the percentage routed to manual review. Next-token accuracy is not a factuality measure. Compare against the extractive baseline at equal review time and count independent incidents rather than individual summary sentences.
Replay serving constraints
At serving, restrict source records to events available by the requested cutoff. Test empty histories, duplicate events, truncated long histories and updates that arrive while a summary is generated. Measure source encoding, token-by-token decoding, evidence validation and p95 total time. Keep output tokens, source IDs and model revision in an audit record, without exposing private event data beyond the authorized view.
Release with a safe fallback
Require zero unsupported severe claims in the reviewed gate set, acceptable severe-fact recall, a bounded review workload and response latency. If evidence is missing or validation fails, show the extractive event list or an unavailable status. Keep the prior workflow as rollback. Any later change to source schema, vocabulary or stop rule requires a new parity run.
Implementation
import torch
from torch import nn
from torch.nn import functional as functional
torch.manual_seed(47)
vocabulary_size = 101
source_ids = torch.tensor([[4, 7, 9, 0], [3, 8, 0, 0]])
target_ids = torch.tensor([[47, 12, 0], [47, 23, 0]])
source_padding = source_ids.eq(0)
target_padding = target_ids.eq(0)
embedding = nn.Embedding(vocabulary_size, 8)
seq2seq = nn.Transformer(d_model=8, nhead=2, num_encoder_layers=1,
num_decoder_layers=1, dim_feedforward=16,
dropout=0.0, batch_first=True)
causal_mask = torch.triu(torch.ones(3, 3, dtype=torch.bool), diagonal=1)
hidden = seq2seq(embedding(source_ids), embedding(target_ids),
tgt_mask=causal_mask, src_key_padding_mask=source_padding,
tgt_key_padding_mask=target_padding,
memory_key_padding_mask=source_padding)
vocabulary_head = nn.Linear(8, vocabulary_size)
logits = vocabulary_head(hidden)
next_token_labels = torch.tensor([[12, 58, 0], [23, 58, 0]])
loss = functional.cross_entropy(logits.reshape(-1, vocabulary_size),
next_token_labels.reshape(-1), ignore_index=0)
loss.backward()
assert logits.shape == (2, 3, vocabulary_size) and torch.isfinite(loss)Performance and operating cost
For B incidents, source length S, target length T and hidden width D, attention score work includes O(BS squared D), O(BT squared D) and O(BTSD). Full sequence decoding adds repeated steps and cache memory, while source tokenization and evidence checks add operational cost. The synthetic forward checks masks and gradient flow only; it does not train on real incidents or measure claim support.
Common Mistakes
- Do not count post-cutoff events as available source evidence.
- Do not claim grounded summaries from next-token loss alone.
- Do not release without a fallback when the source ledger is incomplete.
