A decoder query may attend to encoded source positions, but padding and future-target access have different mask rules.
Encoder-decoder cross-attention and source masks
Separate three attention paths
The encoder can read the available source sequence in both directions. Decoder self-attention may read only earlier target positions when generating autoregressively. Cross-attention uses decoder states as queries and encoded source states as keys and values. A source key-padding mask blocks meaningless padded source positions. It is not a causal mask on the source. The attention-mask lesson defines the shape and polarity checks for each path.
Keep source length physical
An incident summary may draw from a timestamped service-event sequence. Source positions should represent events known before the summary request, not post-resolution notes. Store source event IDs, cutoff time, tokenization revision and padding lengths. A model that reads events after the alert appears to summarize well but cannot do so when deployed. Group train and final test by incident, since two summaries of one incident are not independent observations.
Check the mask at the boundary
For a batch of two incident records, source lengths of five and three require a mask that hides only the last two positions of the second record. A Boolean key-padding mask commonly uses True for blocked positions; verify the chosen API rather than guessing. The code checks that the second query places zero attention weight on padded source positions. Inspect finite outputs when one source is empty; many implementations require an explicit no-source route.
Preserve source evidence
Attention weights are routing coefficients, not proof that generated text is supported. Keep a link from each generated statement to available source events and compare against a simple extractive summary. A long source can dilute rare severe events under truncation; record what was cut and test incident severity slices. The applied project rejects summaries that invent or omit material facts.
Profile both sequence axes
Cross-attention compares each target position with source positions. Long incident histories and long generated summaries therefore raise both compute and memory. Cache the encoded source when decoding many target tokens, but re-encode when the event ledger changes. Measure p95 tokenization, encoding, incremental decoding and evidence validation rather than only one attention layer.
Implementation
import torch
from torch import nn
torch.manual_seed(47)
encoded_events = torch.rand(2, 5, 8)
decoder_queries = torch.rand(2, 3, 8)
source_padding = torch.tensor([[False, False, False, False, False],
[False, False, False, True, True]])
cross_attention = nn.MultiheadAttention(8, 2, batch_first=True)
context, weights = cross_attention(decoder_queries, encoded_events,
encoded_events, key_padding_mask=source_padding,
need_weights=True)
assert context.shape == (2, 3, 8)
assert weights.shape == (2, 3, 5)
assert torch.allclose(weights[1, :, 3:], torch.zeros_like(weights[1, :, 3:]))
assert torch.isfinite(context).all()Performance and operating cost
For batch B, target length T, source length S and width D, cross-attention score work is roughly O(BTSD), with O(BTS) attention-weight storage when weights are materialized. Projection costs add terms proportional to sequence length and D squared. Source encoding and target self-attention add separate work. The example verifies mask geometry only; it does not establish grounded incident summaries or a fast decoder.
Common Mistakes
- Do not apply a target causal mask as if it were a source padding mask.
- Do not include post-decision events in an input for a live summary.
- Do not treat an attention weight as proof that a statement is supported.
