Skip to content
AITroveRead. Build. Understand.
Make this comfortable

CTC greedy collapse and transcription error audit

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Greedy CTC decoding merges consecutive identical frame IDs before removing blanks; evaluation then has to measure the fields people actually read.

Collapse in the correct order

Take the most likely class at each encoded time step. First merge adjacent repeated IDs, then remove blank IDs. A blank between identical character IDs separates two emitted characters. Removing blanks first would merge them and turn a legitimate double letter into one. The pure-Python decoder tests that boundary. Preserve the chosen alphabet map and text-normalization policy with the model; a character index has no meaning without its revision. The alignment lesson establishes blank and length constraints.

Know what greedy decoding misses

The most likely path does not always collapse to the most probable output sequence, because many paths may support one transcript. Greedy decoding is cheap and deterministic, so it is a useful baseline. If errors on ambiguous line items matter, test beam search with a fixed beam width and document whether it uses a language model or field grammar. Compare accuracy, memory and p95 decoding time. A larger beam cannot repair a line crop that omitted the first digit.

Score edits with a clear denominator

Character error rate counts substitutions, deletions and insertions divided by reference characters. A line with no characters needs an explicit policy rather than division by zero. Compute error per line and aggregate edit counts across lines for a micro rate; the mean of per-line rates answers a different question and can overweight tiny strings. Field accuracy for totals, dates and identifiers is often more consequential than a small global character-rate change. Keep punctuation and spacing normalization fixed before comparing models.

Inspect systematic blank behavior

Track blank-frame fraction and output-length distribution, then review empty predictions and repeated-character deletions. A model can achieve a finite training loss and still output mostly blanks early in training or on out-of-domain fonts. Stratify by line width, contrast, blur and capture device. Do not lower the reported error by discarding low-confidence lines; show coverage and risk together. Abstention policy belongs in the final workflow.

Audit reading order and release thresholds

A line recognizer assumes that the crop contains one sequence in reading order. Multi-column receipts require a line detector and ordering rule before CTC decoding. Review the highest-cost fields and manual-correction time, not only raw character error. Freeze normalization, decoder and validation thresholds before opening the final held-out receipts. The project gates the recognizer on field-level quality and latency.

Implementation

python
def greedy_ctc(frame_ids, blank_id=0):
    collapsed = []
    previous = None
    for class_id in frame_ids:
        if class_id != previous:
            collapsed.append(class_id)
        previous = class_id
    return [class_id for class_id in collapsed if class_id != blank_id]

frame_classes = [0, 1, 1, 0, 2, 0, 2, 2]
assert greedy_ctc(frame_classes) == [1, 2, 2]
assert greedy_ctc([0, 0, 0]) == []
alphabet = {1: "S", 2: "7"}
recognized_line = "".join(alphabet[class_id]
                          for class_id in greedy_ctc(frame_classes))
assert recognized_line == "S77"

Performance and operating cost

Greedy decoding is O(T) time and O(T) output memory for T frame IDs, or O(U) output memory if only emitted symbols are retained. A beam search retains multiple partial prefixes and costs more time and memory as beam width grows. Character error requires an edit-distance calculation, commonly O(RP) time for reference length R and prediction length P, with row-wise O(min(R, P)) memory. Report field correction cost and decoder latency at the same width and batch policy used in serving.

Common Mistakes

  • Do not remove blanks before collapsing adjacent repeats.
  • Do not compare character error after changing text normalization between models.
  • Do not present a greedy path as necessarily the highest-probability transcript.

Read next

ai-data
deep-learning
Storage details