Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Project: distill and release a receipt-quality student

Last updated: 6 Oct 20265 min read
project
AdvancedBy AITrove Editorial

Transfer a reviewed teacher into a smaller receipt classifier, then prove that the student meets class-specific decision and device limits before release.

Pin teacher, student and populations

Freeze a teacher trained only on the allowed training receipts and verify its class map, preprocessing and held-out defect behavior. Choose a smaller student architecture before tuning distillation weights. Partition physical receipts into training, development, calibration and final test identities; do not reuse the final test to select temperature or channels. Train a hard-label-only student as a fair baseline. The soft-target contract fixes the teacher-to-student objective.

Run controlled ablations

Use the same student initialization policy, training examples and update budget for hard labels alone, hard plus teacher logits, and an optional feature-hint variant. Freeze teacher parameters and keep it in evaluation mode for every run. Save class-map and preprocessing hashes with cached logits if caching is used. Check that training loss reduction is not merely copying teacher errors on clipped receipts. Feature alignment adds an adapter that must not sneak into the exported classifier.

Measure decisions by slice

On development data compare quality-class recall, false accepts and teacher-student disagreement, including low light, blur, clipped edges and capture devices. Fit any serving calibration on a separate calibration group and choose a manual-review threshold before opening final test. The code below counts disagreements and errors for a fixed synthetic audit batch; it is a metric-definition check rather than evidence that a real student passed. Selective risk defines the review tradeoff.

Test the target runtime

Export the student without training-only teacher and adapter layers. Reload it in a clean process and compare fixed-batch logits with its pre-export graph. Measure cold-start, resident memory, median and tail latency with preprocessing at the intended device batch size. A narrow student may still be slower if its operators lack efficient kernels. Compare a later quantized student as a separate candidate, calibrating it on representative data and rerunning the same final gates. The edge release supplies the device gate pattern.

Make failure visible

Set acceptable rare-defect recall loss, false-accept count, review workload, memory and 95th-percentile latency before final test. If the distilled student fails one gate, keep the teacher or hard-label student as appropriate; do not compensate by reporting only average accuracy. Deliver the ablation table, disagreement examples, class-map manifest, calibration fit, exported-graph parity and rollback bundle. A smaller model is an engineering option, not a reason to hide changed decisions from operators.

Implementation

python
import torch

teacher_logits = torch.tensor([[3.1, 0.4, -0.2], [0.2, 0.8, 2.1],
                               [0.1, 2.5, 0.3], [1.3, 0.9, 0.4],
                               [0.4, 0.6, 1.8], [2.0, 0.5, 0.3]])
student_logits = torch.tensor([[2.7, 0.5, 0.1], [0.2, 1.4, 1.3],
                               [0.3, 2.0, 0.4], [1.1, 1.3, 0.5],
                               [0.3, 0.6, 1.4], [1.8, 0.7, 0.2]])
reviewed_labels = torch.tensor([0, 2, 1, 0, 2, 0])
teacher_decisions = teacher_logits.argmax(dim=1)
student_decisions = student_logits.argmax(dim=1)
disagreement = teacher_decisions.ne(student_decisions)
student_errors = student_decisions.ne(reviewed_labels)
teacher_errors = teacher_decisions.ne(reviewed_labels)
audit = {"teacher_errors": int(teacher_errors.sum()),
         "student_errors": int(student_errors.sum()),
         "disagreements": int(disagreement.sum())}
assert audit == {"teacher_errors": 0, "student_errors": 2, "disagreements": 2}

Performance and operating cost

Distillation pays for teacher inference during student training but deploys only the student. Feature-hint experiments add adapter and teacher-feature storage, while calibration and device benchmarking add separate work. The six-record code audit is O(NC) for N records and C classes and cannot estimate population error. Real release tests need adequate counts in rare slices and repeated tail-latency measurements. Measure end-to-end device cost; a parameter ratio is not a latency guarantee.

Common Mistakes

  • Do not select distillation settings on final-test receipts.
  • Do not package the teacher or feature adapter into a student-only latency claim.
  • Do not hide a clipped-edge false accept behind a lower aggregate error count.

Read next

Continue the workflow: Project: adapt a service-event classifier with low-rank weights.

ai-data
deep-learning
Storage details