Train a line-level recognizer on transcript-only labels, then prove that encoded widths, blank collapse, field accuracy and review latency hold on unseen receipts.
Project: recognize receipt line items with CTC
Create line and transcript contracts
Extract one reading-order line per image and keep the parent physical receipt ID. Normalize orientation, grayscale range and transcript characters under a versioned rule; retain the raw image and reviewed text. Split by physical receipt before crops or synthetic distortions. Record blank ID zero and a fixed alphabet. The miniature model below takes already encoded time-step features and performs one valid loss update; it does not claim to detect lines or read photographs. The CTC contract defines legal lengths.
Audit width and supervision
For every batch, derive encoded time lengths from the actual feature extractor output and original valid crop width, not the padded batch width. Reject empty transcripts unless the workflow explicitly models them. Flag unknown characters and target repetitions that cannot fit the available frames. Save the excluded-line report so quality metrics cannot quietly improve by dropping hard cases. The code provides seven and nine valid time steps for two transcripts and checks that labels contain no blank.
Train and decode separately
Monitor finite loss, gradient norms and output blank fraction. Decode with the frozen greedy collapse rule; only add beam search if it improves held-out field accuracy enough to justify its time cost. Report character edits, exact line matches, exact total/date matches and manual correction minutes by store format and camera. The decoder lesson defines repeated-character behavior. Compare with a non-neural OCR baseline under identical crops and normalization.
Validate the pipeline rather than one tensor
Run line detection, ordering, crop normalization, encoder, decoder and field parsing on held-out full receipts. Check that a total on the neighboring line is not attributed to the wrong item. Group uncertainty by receipt ID because several line images from one receipt are correlated. Measure p95 wall time including cropping and decoding, and inspect empty or truncated outputs. If a crop omits half a line, a better CTC model cannot recover its missing content.
Release with a fallback
Predeclare minimum exact-field accuracy, maximum deletion rate on repeated characters, review capacity and p95 end-to-end latency. Package alphabet, blank ID, crop policy, encoded-length calculation, model, decoder and parser versions. Re-run a fixed receipt set after upgrades; route low-confidence or malformed fields to review while retaining the previous OCR path. The deliverable is a field-level error ledger and reproducible deployment bundle, not just a declining training-loss curve.
Implementation
import torch
from torch import nn
torch.manual_seed(47)
line_features = torch.randn(9, 2, 6)
valid_time = torch.tensor([9, 7], dtype=torch.long)
transcript_ids = torch.tensor([1, 2, 2, 3, 1], dtype=torch.long)
transcript_sizes = torch.tensor([3, 2], dtype=torch.long)
assert valid_time.max().item() <= line_features.size(0)
assert transcript_sizes.sum().item() == transcript_ids.numel()
assert not (transcript_ids == 0).any()
character_head = nn.Linear(6, 5)
optimizer = torch.optim.AdamW(character_head.parameters(), lr=0.0007)
optimizer.zero_grad(set_to_none=True)
log_probs = character_head(line_features).log_softmax(-1)
loss = nn.CTCLoss(blank=0, zero_infinity=False)(
log_probs, transcript_ids, valid_time, transcript_sizes)
assert torch.isfinite(loss)
loss.backward()
optimizer.step()Performance and operating cost
A line encoder generally dominates inference cost, while CTC scoring must process valid encoded time steps and transcript lengths. This tiny linear head has O(TBCF) multiply cost for T steps, B lines, C classes and F input features, plus O(TBC) logits memory; a real convolutional encoder adds its own activations. Padding every crop to the widest batch item wastes compute, so bucket similar widths while retaining true lengths. Evaluate end-to-end receipt latency, since detection and ordering may cost more than the character head.
Common Mistakes
- Do not split line crops from one physical receipt across train and test.
- Do not use padded width as every sample’s CTC input length.
- Do not optimize character error while ignoring totals, dates and manual correction cost.
