Feature extraction and fine-tuning use the same pretrained weights but change which state is allowed to move, how much memory training needs and what can overfit.
Frozen features versus fine-tuning a pretrained backbone
Establish a frozen baseline
Replace the pretrained classifier with a new head for the target labels, freeze every backbone parameter and train only the head. Keep the backbone in evaluation mode if its normalization buffers are intended to remain fixed. Match the preprocessing expected by the selected weights; changing color order or normalization can make the features meaningless even though tensor shapes match. Receipt-quality labels provide a concrete target while the pretrained backbone supplies initial visual features.
Unfreeze only when evidence warrants it
A frozen baseline is fast to run and makes a useful lower-risk comparison. If errors cluster on domain-specific marks, unfreeze a late block and measure whether held-out groups improve. New head parameters start from a different distribution than pretrained features, so use a separately controlled learning rate for each parameter group. Training every layer from the first update on a small dataset can move useful features before the head stabilizes. Validation discipline decides whether adaptation helped.
Separate gradients from module mode
requires_grad false prevents gradient accumulation for frozen parameters, but train mode can still update BatchNorm buffers and apply dropout. Use a wrapper or training-loop rule that reapplies eval to a frozen backbone after calling model.train on the full model. If a late block is unfrozen, choose explicitly whether its normalization buffers also adapt. Do not infer the choice from optimizer parameter groups. Normalization state belongs in the checkpoint and test plan.
Budget the two phases
With a frozen backbone, no backward graph through its internal layers is needed; only the new head requires gradient storage. Fine-tuning late layers stores their activations for backward, increasing memory and training time. Keep frozen-feature caching only when preprocessing is deterministic and the backbone truly fixed; random augmentation changes features each epoch. If cached features are derived from images with sensitive information, apply the same retention and access controls as the original inputs.
Keep comparison fair
Use the same identity-disjoint split, label map, target metric and decision threshold policy for frozen and fine-tuned models. Compare by source device and defect type, not only the mean score. A small average gain can hide a serious regression on low-contrast receipts. Save both candidates, preprocessing configuration and validation report; choose with a stated rule before opening the final test population.
Implementation
import torch
from torch import nn
from torchvision.models import resnet18, ResNet18_Weights
selected_weights = ResNet18_Weights.IMAGENET1K_V1
quality_backbone = resnet18(weights=selected_weights)
input_features = quality_backbone.fc.in_features
quality_backbone.fc = nn.Linear(input_features, 3)
for parameter_name, parameter in quality_backbone.named_parameters():
parameter.requires_grad = parameter_name.startswith("fc.")
quality_backbone.train()
for module_name, module in quality_backbone.named_modules():
if module_name and not module_name.startswith("fc"):
module.eval()
head_optimizer = torch.optim.AdamW(quality_backbone.fc.parameters(), lr=0.00037)
assert all(not parameter.requires_grad for parameter in quality_backbone.layer1.parameters())
assert quality_backbone.bn1.training is False
assert quality_backbone.fc.training is True
preprocess_for_inference = selected_weights.transforms()Performance and operating cost
Feature extraction still runs a backbone forward pass, but saves much less backward state and updates only the head. Fine-tuning a late stage adds its parameter gradients, optimizer state and retained activations. Full fine-tuning can approach the cost of training the whole network. Pretrained weights may download once, so pin their identity and cache the artifact in a controlled build rather than depending on a network fetch at serving time.
Common Mistakes
- Do not freeze parameters while leaving running normalization buffers free to drift by accident.
- Do not compare models built with different image normalization or label splits.
- Do not deploy a model that downloads its pretrained weights during a request.
Read next
- Residual blocks and normalization state
- Training and validation modes: measure the model you will serve
- Project: classify receipt image quality with a checked training contract
- Image augmentation: split originals first and preserve the label
- Staged unfreezing and domain-shift audits
Continue the workflow: Structured pruning and latency parity.
Continue the workflow: Project: pretrain receipt features from unlabeled captures.
Continue the workflow: Vision position grids, resolution changes and transfer.
Continue the workflow: LoRA rank, scaling and frozen-base contracts.
