Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Decoding stop rules, sampling and beam-score accounting

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Token generation is a decision policy over logits; stopping, forbidden tokens and sequence scores affect output as much as the next-token model does.

Mask invalid tokens before sampling

A decoder may forbid padding, a start marker after the prefix or application-specific events. Apply those masks to raw logits before softmax, and reject a step where every candidate is forbidden. Renormalize the remaining distribution. For top-p sampling, sort valid probabilities and keep the smallest prefix whose cumulative mass reaches the threshold, including the token that crosses it. Sample from the retained, renormalized probabilities with an explicit random generator. The code demonstrates one step with a fixed seed; it is not a language-quality claim.

Define stopping independently of padding

An end token terminates a generated sequence; padding only fills a batch tensor and must not be interpreted as a model-selected end. Set a maximum new-token count as a hard guard even when the model usually emits an end token. State whether the end token is included in stored output and whether a minimum generated length suppresses it temporarily. If a model never emits end, return a capped status rather than silently presenting a complete sequence. The shifted training target must include the intended end event.

Score beams with the same units

Beam search accumulates log probabilities for each candidate sequence and keeps a bounded number of live hypotheses. Raw summed log probability favors short sequences because every added token contributes a nonpositive term. A length normalization or penalty changes the ranking; document its formula and when it is applied. Finished beams must not keep collecting token probability. When live beams are reordered, reorder their per-layer caches and any batch-specific masks with the same parent indices. Cache parity is part of beam correctness.

Separate diversity from correctness

Greedy decoding is deterministic under a fixed model and preprocessing. Sampling introduces variation; beam search optimizes a score under a finite search width, not factual or task correctness. Compare output validity, duplicate sequences, event-order constraints and error costs under a fixed held-out prefix set. Top-p and temperature change the distribution; a lower temperature is not a substitute for a trained model that respects domain rules. For service events, invalid order may need a constrained decoder or a post-decode rejection path.

Budget the requested length

Latency rises with generated tokens, and beam search adds model and cache work for multiple hypotheses. Report time to first event, later-event rate, p95 whole-request latency and capped-output fraction. A model with good per-token accuracy can be unusable if it rarely emits end. Keep decoding settings in the serving artifact, not in an undocumented caller default. The applied project compares policies under event-sequence contracts.

Implementation

python
import torch
from torch.nn import functional as functional

torch.manual_seed(47)
event_logits = torch.tensor([1.2, 0.8, 2.1, -0.3, 0.4])
start_token, padding_token = 0, 4
masked_logits = event_logits.clone()
masked_logits[[start_token, padding_token]] = float("-inf")
if not torch.isfinite(masked_logits).any():
    raise ValueError("no permitted next event")
sorted_logits, sorted_ids = torch.sort(masked_logits, descending=True)
sorted_probabilities = functional.softmax(sorted_logits, dim=0)
cumulative_mass = torch.cumsum(sorted_probabilities, dim=0)
keep = cumulative_mass - sorted_probabilities < 0.82
candidate_ids = sorted_ids[keep]
candidate_probabilities = sorted_probabilities[keep]
candidate_probabilities /= candidate_probabilities.sum()
sampled_offset = torch.multinomial(candidate_probabilities, 1)
next_event = int(candidate_ids[sampled_offset])
assert next_event not in (start_token, padding_token)
assert len(candidate_ids) >= 1

Performance and operating cost

Sorting V vocabulary logits for top-p selection costs O(V log V) time per generated token; selecting a fixed top-k can be cheaper. Beam search generally multiplies decode-state storage and candidate scoring by beam width, though batching can improve hardware use. A maximum of K new tokens bounds work and cache growth for a request. The five-logit code is a policy check; real latency includes model forward passes, cache movement and output serialization, which should be measured at realistic prefix and output lengths.

Common Mistakes

  • Do not treat padding as a generated end token.
  • Do not discard the token that first crosses the top-p cumulative threshold.
  • Do not compare beam sums across lengths without the stated normalization rule.

Read next

Continue the workflow: CTC greedy collapse and transcription error audit.

Continue the workflow: Project: generate grounded summaries of service incidents.

ai-data
deep-learning
Storage details