Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Teacher forcing, target shifts and generation parity

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A sequence-to-sequence model trains against known previous tokens but must generate from its own history at serving time.

Shift the target once

If the reviewed incident summary is a token sequence ending in an end marker, the decoder input starts with a beginning marker followed by every target token except the final one. Labels are the original target tokens. Never feed the token being predicted into the same input position. The code constructs the input and label arrays and checks the offset. The causal shift lesson covers the same off-by-one failure for decoder-only training.

Mask target padding separately

A batched summary can end after five tokens in one incident and twelve in another. Ignore padding labels in the loss while still preventing decoder self-attention from reading future target tokens. Source padding belongs to cross-attention. Track beginning, end and pad token IDs as part of the vocabulary revision; mixing them can create an apparently good loss and endless generation. Check the empty-summary and maximum-length cases explicitly.

Measure exposure at generation

During teacher-forced training, every previous target token is correct. At serving, an early mistaken word becomes part of subsequent context and can compound. Evaluate by actually decoding complete summaries, not only by next-token loss. Compare greedy and beam decoding with a fixed length rule, and inspect repeated phrases, premature end markers and missing severe facts. The decoding lesson defines a stable stop contract.

Tie claims to available events

A service summary should identify which source event supports each critical claim. Record the incident cutoff, source event IDs and generated tokens, then let reviewers mark unsupported statements. An extractive baseline may be safer for a sparse event ledger. If a source is missing, use a declared unavailable response rather than completing a plausible story from learned patterns. The model must not silently fill gaps with invented operational details.

Keep release parity measurable

Replay the same incident through training preprocessing and serving preprocessing, checking token IDs, masks and logits at the first generated step. Then perform full decoding with only past generated tokens. Measure factual error, severe-event omission, p95 latency and fallback frequency by incident type. The project gates these outcomes before any operator-facing release.

Implementation

python
START_ID, END_ID, PAD_ID = 47, 58, 69

def shifted_summary(reviewed_tokens):
    if not reviewed_tokens or reviewed_tokens[-1] != END_ID:
        raise ValueError("reviewed summary must end with END_ID")
    if PAD_ID in reviewed_tokens or START_ID in reviewed_tokens:
        raise ValueError("special tokens are not summary content")
    decoder_input = [START_ID] + reviewed_tokens[:-1]
    labels = list(reviewed_tokens)
    assert len(decoder_input) == len(labels)
    return decoder_input, labels

decoder_input, labels = shifted_summary([72, 83, 94, END_ID])
assert decoder_input == [START_ID, 72, 83, 94]
assert labels == [72, 83, 94, END_ID]
assert all(decoder_input[position] != labels[position]
           for position in range(len(labels)))

Performance and operating cost

Constructing shifted labels uses O(T) time and memory for T target tokens. Teacher-forced training can score all target positions in parallel within a batch, while autoregressive serving emits tokens one step at a time and can incur T model invocations. Caching past keys and values reduces repeated decoder work but uses sequence-growing memory. The snippet checks only token alignment; it does not validate a tokenizer or full text quality.

Common Mistakes

  • Do not expose the next target token at its own prediction position.
  • Do not evaluate only teacher-forced loss when serving uses free generation.
  • Do not treat a missing source event as permission to invent one.

Read next

ai-data
deep-learning
Storage details