Distillation trains a smaller student against task labels and softened teacher predictions, then tests whether the student preserves target-task decisions at lower serving cost.
Teacher–student distillation objective
Keep the teacher frozen
A large parcel-image teacher can supply class logits on the target training cohort. A smaller student learns from true damage labels and the teacher’s relative scores; the teacher is not updated by the student loss. The code computes one hard-label cross-entropy and one temperature-scaled teacher-to-student divergence, then combines them. It is an objective calculation, not a full neural training framework. Logit and loss mechanics give the optimization context.
Use the temperature consistently
Divide both teacher and student logits by the same positive temperature before softmax. A higher value exposes more of the relative mass in lower-ranked classes. The conventional temperature-squared factor keeps the soft-loss scale from vanishing as the distribution softens; its coefficient still requires development tuning. Do not copy a teacher’s incorrect confidence as ground truth. Calibration is a separate check.
Protect the target holdout
Generate teacher targets for training cases only. Select student size, loss mixture and checkpoint on development data, then evaluate on a target holdout untouched by those choices. If the teacher was trained on target-test images, a favorable student comparison can inherit overlap. The provenance audit records what is known about teacher training.
Compare with a student trained without the teacher
The relevant gain is not only whether the student approaches teacher accuracy. Compare a same-size student trained on true labels alone, using identical target splits and serving hardware. Report class-specific misses, especially expensive damage types, and the full-pipeline speed. Device profiling defines the cost side.
Know the limits of imitation
Teacher predictions may carry source-domain shortcuts or annotation mistakes into the student. A small model may be less accurate on rare camera conditions even when mean loss improves. Inspect paired target errors and retain a manual-review fallback. The release project treats that as a decision condition.
Implementation
from math import exp, log
teacher_logits = (2.4, 0.9, -0.7)
student_logits = (1.8, 1.1, -0.4)
true_class = 0 # Parcel requires seal review.
temperature = 2.5
soft_weight = 0.6
def softmax(logits, scale=1.0):
if scale <= 0:
raise ValueError("positive temperature required")
shifted = [value / scale for value in logits]
peak = max(shifted)
exponentials = [exp(value - peak) for value in shifted]
total = sum(exponentials)
return [value / total for value in exponentials]
teacher_soft = softmax(teacher_logits, temperature)
student_soft = softmax(student_logits, temperature)
soft_loss = sum(teacher_share * log(teacher_share / student_share)
for teacher_share, student_share in zip(teacher_soft, student_soft))
soft_loss *= temperature ** 2
hard_loss = -log(softmax(student_logits)[true_class])
combined_loss = soft_weight * soft_loss + (1 - soft_weight) * hard_loss
assert 0 < soft_loss < hard_loss
assert combined_loss > 0Performance and operating cost
For C output classes, this per-case objective costs O(C) time and O(C) temporary probabilities. Training the student also requires its forward and backward passes plus teacher inference or stored teacher logits. At serving time only the student should run; measure the complete pipeline rather than assuming parameter reduction guarantees speed.
Common Mistakes
- Do not soften only the teacher logits while comparing against an unsoftened student distribution.
- Do not train or select from an untouched target holdout.
- Do not accept average imitation quality without rare-class and device checks.
Read next
- Model inference budget and device profile
- Structured pruning and compute shape
- Compression Pareto review and shadow check
- Model compression release project
Continue the workflow: Multi-task head losses with observed-label masks.
