Matching internal teacher features can guide a narrow student, but dimensions, spatial resolution and representation scale must be reconciled deliberately.
Teacher feature adapters and layer alignment
Choose an internal comparison point
A teacher and student need not have matching depth or hidden width. Select a teacher feature after a stable block and a student feature with related receptive field or semantic role. For image networks, an early edge detector should not be forced to match a late global decision representation without a reason. Record layer names, pre- or post-activation choice, spatial shape and normalization behavior. Freeze the teacher and capture features under the same input preprocessing. Convolution geometry explains why stride changes spatial support.
Map shapes with a learned adapter
If student features have twelve channels and teacher features have twenty-four, use a small projection such as a linear layer or one-by-one convolution to map the student feature into the teacher space. Match spatial resolution with a documented pooling or interpolation rule when needed. Train the adapter with the student and exclude it from final inference if it is used only for the auxiliary loss. A shape-compatible projection does not prove semantic alignment; compare downstream decisions and feature variance. The code uses vector features so the shape contract is visible.
Normalize the auxiliary objective
Raw mean squared error can be dominated by teacher feature magnitude. Normalize features or choose an explicit weighting so the auxiliary term does not swamp human-label loss. Report each loss separately and inspect gradient norms for student backbone, adapter and classifier. A near-zero feature loss might mean the adapter learned a trivial mapping or the teacher features had low variance. Use a feature collapse check and compare a logit-only distillation baseline. Soft-target learning is a simpler reference.
Handle residual and spatial differences
For convolutional teachers, feature width may vary along with height and width. Interpolation can blur local defects while pooling can discard them; neither is neutral. Prefer comparison points with similar effective receptive fields, and inspect cut-off edges after any resize. If a teacher has residual connections or normalization running state, extraction must come from the intended post-block tensor. Verify hooks are removed after use and do not retain tensors across batches, which can leak memory in long training runs.
Test whether the hint helps
Train equal-budget students with hard labels alone, hard plus soft logits, and hard plus soft plus feature alignment. Select objective weights on a development split, then report a held-out defect table and device latency. The adapter can be thrown away at serving time only if the exported student output does not depend on it. Keep a fixed-input parity check for the exported graph. A teacher hint that raises average accuracy but degrades an important rare defect is not a successful transfer. The student project uses this ablation.
Implementation
import torch
from torch import nn
from torch.nn import functional as functional
torch.manual_seed(47)
receipt_features = torch.randn(5, 9)
teacher_backbone = nn.Sequential(nn.Linear(9, 24), nn.ReLU())
student_backbone = nn.Sequential(nn.Linear(9, 12), nn.ReLU())
feature_adapter = nn.Linear(12, 24)
for parameter in teacher_backbone.parameters():
parameter.requires_grad_(False)
teacher_backbone.eval()
with torch.no_grad():
teacher_features = teacher_backbone(receipt_features)
student_features = student_backbone(receipt_features)
aligned_student = feature_adapter(student_features)
hint_loss = functional.mse_loss(functional.normalize(aligned_student, dim=1),
functional.normalize(teacher_features, dim=1))
hint_loss.backward()
assert aligned_student.shape == teacher_features.shape == (5, 24)
assert all(parameter.grad is None for parameter in teacher_backbone.parameters())
assert torch.isfinite(hint_loss)Performance and operating cost
The adapter adds O(Bdsdt) multiply work for batch B, student width ds and teacher width dt in this dense example, as well as a teacher forward during training. A spatial adapter has additional height and width factors. Feature storage and hooks can raise training memory substantially if several layers are compared. If the adapter serves only the loss, it can be omitted from the final student graph; verify that omission in a clean export. Extra auxiliary compute is justified only by measured held-out gains or a better release tradeoff.
Common Mistakes
- Do not align unrelated layer positions solely because their tensors can be resized.
- Do not allow adapter output magnitude to dominate the hard-label objective.
- Do not claim a smaller serving model until the adapter is absent from the exported graph.
