Skip to content
AITroveRead. Build. Understand.
Make this comfortable

CTC time lengths, blank labels and alignment feasibility

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Connectionist temporal classification trains a sequence recognizer without per-frame character labels, but every encoded length and repeated symbol must fit a valid alignment.

State the alignment problem

A receipt line image may produce thirty visual time steps while its transcript contains eleven characters. No annotator has marked which column owns each character. CTC sums the probabilities of frame-level paths that collapse to the transcript. The output vocabulary includes a blank symbol that is not a transcript character. This is useful when monotonic left-to-right alignment is plausible; it is a poor fit for reading order that jumps between multiple columns without a separate layout policy. Pixel masks solve a different supervision problem.

Calculate encoded length after the encoder

The input length passed to the loss is the number of valid encoder time steps, not raw image width, original audio samples or the length of a greedy-decoded transcript. Every stride, pooling and crop changes that length. If a batch pads images to a common width, retain the pre-padding encoded length for each item. Check all input lengths against the logits time dimension. The loss convention used here expects log probabilities arranged as time, batch, classes; silently swapping time and batch can produce plausible tensor values with the wrong semantics.

Reserve blank and validate targets

Choose one blank ID, such as zero, and keep it out of the target transcript. A repeated character needs a blank or another frame-level separation to survive collapse: a transcript of two consecutive identical characters can require more time steps than its raw target length. The code tests a repeated-character target of length three against seven encoder steps. Check the alphabet map, normalized transcript and lengths before the loss. An unknown character should be rejected or mapped by a documented policy, not silently dropped.

Treat impossible alignments as data failures

If an image is too narrow after the encoder, the transcript has no legal CTC path and loss may become infinite. Some implementations can zero such losses, but that can hide a systematic preprocessing error. Count and inspect rejected examples, especially long line items and repeated-character cases. Fix encoder resolution or data preparation rather than quietly training on a smaller population. Keep target-length distribution and the ratio of encoded steps to characters in the dataset report.

Separate loss from transcription quality

A finite CTC loss proves only that the alignment calculation ran. Decode predictions with a fixed blank-collapse rule, then report character and field error by receipt type and capture device. Track output blank rate, deletion-heavy failures and line-order mistakes. The decoder lesson prevents the common error of removing blanks before merging repeated IDs. The project joins training with field-level review.

Implementation

python
import torch
from torch import nn

torch.manual_seed(47)
encoder_logits = torch.randn(7, 1, 5, requires_grad=True)
line_transcript = torch.tensor([1, 2, 2], dtype=torch.long)
encoded_lengths = torch.tensor([7], dtype=torch.long)
transcript_lengths = torch.tensor([3], dtype=torch.long)
assert encoded_lengths.max().item() <= encoder_logits.size(0)
assert not (line_transcript == 0).any()
ctc = nn.CTCLoss(blank=0, reduction="mean", zero_infinity=False)
loss = ctc(encoder_logits.log_softmax(-1), line_transcript,
           encoded_lengths, transcript_lengths)
assert torch.isfinite(loss)
loss.backward()
assert torch.isfinite(encoder_logits.grad).all()

Performance and operating cost

For T encoded steps, U transcript symbols and C output classes, logits alone cost O(TC) memory per example; the alignment calculation adds work that grows with both T and U, often described as O(TU) state transitions for a fixed sequence. A wider image increases encoder activations and CTC work even when its transcript remains short. The example is an alignment and gradient check, not a trained OCR system. Measure accepted versus rejected lines and actual throughput for the target widths.

Common Mistakes

  • Do not pass raw image width as encoder input length after strided convolutions.
  • Do not put the blank ID in a transcript or pad with an ID counted as a target.
  • Do not hide impossible alignments by zeroing infinite loss without an audit.

Read next

Continue the workflow: Audio sample rates, short-time spectra and window contracts.

ai-data
deep-learning
Storage details