Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Staged fine-tuning and checkpoint selection

Last updated: 5 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Staged fine-tuning first trains a task head with the encoder frozen, then tests a limited encoder update using an untouched validation role and a final target holdout.

Start from an interpretable probe

A frozen encoder with a trained head establishes a cheap baseline. Unfreeze a declared upper block only after that baseline is measured. Changing every layer at once can destroy useful pretrained features, increase memory use and overfit a small target set. The exact layer choice is an experiment, not a universal rule. The probe lesson supplies the starting point.

Use a separate, smaller update rate for reused weights

The head begins near a new task solution; the encoder already contains a learned representation. It is common to test a smaller update rate for reused layers than for the head. Log which parameters are trainable, their update rates, optimizer state and checkpoint. Some layers maintain running statistics even when their weights are frozen, so inspect the training/inference behavior of the chosen framework. Checkpoint state must include more than weights.

Select only on development data

The code compares frozen and partially tuned checkpoint records using predeclared validation loss, then reports a final-test value only after selection. Its records are illustrative; it does not implement a neural optimizer. If several unfreezing depths, augmentations and seeds are tried, the validation set becomes part of a large search. A later independent target cohort is needed to assess the chosen design. Untouched testing states the boundary.

Measure the tradeoff across slices

A tuned encoder may improve familiar photos while worsening images from a new camera or packaging type. Compare paired errors, calibration and review load by slice, with support. Do not promote a small average gain if it hides a costly miss increase at one depot. Negative transfer is an empirical finding on the target task.

Include serving cost in the decision

Fine-tuning does not necessarily make inference slower if architecture stays fixed, but training costs rise and the new checkpoint needs storage, validation and rollback. If it changes input size or adds layers, inference may also change. Serving latency belongs beside target error. The release review compares both.

Implementation

python
# Candidate records are produced by separate training runs, not this selector.
checkpoints = [
    {"id": "frozen-head-v3", "stage": "frozen", "validation_loss": 0.41,
     "trainable_parameters": 513},
    {"id": "upper-block-v1", "stage": "partial", "validation_loss": 0.36,
     "trainable_parameters": 180513},
    {"id": "all-layers-v1", "stage": "full", "validation_loss": 0.44,
     "trainable_parameters": 2200513},
]

def choose_checkpoint(training_records):
    if len({record["id"] for record in training_records}) != len(training_records):
        raise ValueError("duplicate checkpoint ID")
    return min(training_records, key=lambda record: record["validation_loss"])

selected = choose_checkpoint(checkpoints)
assert selected["id"] == "upper-block-v1"
sealed_test_losses = {"frozen-head-v3": 0.43, "upper-block-v1": 0.39,
                      "all-layers-v1": 0.50}
final_test_loss = sealed_test_losses[selected["id"]]  # Revealed after selection.
assert final_test_loss == 0.39

Performance and operating cost

Scanning K checkpoint summaries costs O(K) time and O(1) extra memory. Fine-tuning a trainable encoder is far more expensive than fitting only a small head because gradients and optimizer state cover more parameters. Retaining multiple checkpoints also consumes storage and review time.

Common Mistakes

  • Do not select a checkpoint using final-test loss.
  • Do not assume frozen weights imply frozen normalization statistics.
  • Do not describe a chosen learning-rate schedule as a universal recipe.

Read next

ai-data
machine-learning
Storage details