Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Frozen embeddings and a linear probe

Last updated: 5 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A linear probe trains a small classifier on fixed target embeddings to test whether the pretrained representation separates the new labels without updating the encoder.

Keep the encoder fixed

Encode parcel photos once with a pinned model and preprocessing contract. Train the probe only on target training parcels; use a separate development group to choose regularization and threshold. Because the encoder does not update, the probe isolates how useful its existing features are for the target task. The contract guide pins the representation and target definition.

Make the small head real

The code fits a two-feature logistic head with batch gradient steps and L2 penalty on the weights. It operates on already-extracted embeddings so it runs without a large image library. Production code should use a tested numerical implementation, standardize dimensions using training-only statistics when appropriate, and monitor convergence. The result is a real training loop, not a line-by-line stand-in for full image model training. Log loss explains the objective.

Test the right baseline

A visually similar source task may still encode the wrong signal for torn seals, glare or lighting at the target depot. Compare the probe with an operational heuristic and, if labels permit, a target-only model. Keep the same parcel-grouped split, decision threshold and false-negative cost. Grouped evaluation prevents repeated photos from inflating the result.

Read probe success carefully

A strong probe shows that a simple head can extract useful information from the frozen vectors on this evaluation set. It does not prove every target subgroup is covered, that the representation is causal, or that unfreezing layers would help. Inspect error by camera, packaging material and depot. The target-slice audit compares paired errors and support.

Measure the serving path

Caching embeddings accelerates repeated experiments but the live service still runs the encoder unless vectors already exist at decision time. Report end-to-end latency, memory and the fallback if feature extraction fails. Inference contracts put the head’s small compute in context. Staged tuning is a later comparison, not a prerequisite.

Implementation

python
from math import exp

# Fixed embeddings from a pinned encoder; labels come from target parcels.
training_parcels = [
    ((-1.4, -0.8), 0), ((-1.0, -1.2), 0), ((-0.8, -0.6), 0),
    ((0.7, 0.9), 1), ((1.2, 0.8), 1), ((0.9, 1.3), 1),
]

def train_damage_probe(rows, steps=180, rate=0.25, penalty=0.02):
    weights = [0.0, 0.0]
    bias = 0.0
    for _ in range(steps):
        gradients = [0.0, 0.0]
        bias_gradient = 0.0
        for embedding, damaged in rows:
            logit = sum(coef * value for coef, value in zip(weights, embedding)) + bias
            probability = 1 / (1 + exp(-logit))
            residual = probability - damaged
            gradients[0] += residual * embedding[0]
            gradients[1] += residual * embedding[1]
            bias_gradient += residual
        for position in range(2):
            gradient = gradients[position] / len(rows) + penalty * weights[position]
            weights[position] -= rate * gradient
        bias -= rate * bias_gradient / len(rows)
    return weights, bias

probe_weights, probe_bias = train_damage_probe(training_parcels)
def damage_probability(embedding):
    logit = sum(coef * value for coef, value in zip(probe_weights, embedding))
    return 1 / (1 + exp(-(logit + probe_bias)))

assert damage_probability((1.0, 1.0)) > 0.8
assert damage_probability((-1.0, -1.0)) < 0.2

Performance and operating cost

For N target rows, D embedding dimensions and T full-batch steps, head training costs O(TND) time and O(D) parameter memory; stored embeddings cost O(ND). Encoder extraction may dominate both. Retain grouped target holdouts rather than spending every labeled parcel on the probe.

Common Mistakes

  • Do not fit scaling or the probe on the final target test.
  • Do not report head-only latency as end-to-end image latency.
  • Do not interpret a high training score on six illustrative rows as deployment evidence.

Read next

Continue the workflow: Binary relevance and label dependence.

Continue the workflow: Embedding pairs, identity labels and leakage-safe splits.

Continue the workflow: Representation collapse and frozen linear probes.

ai-data
machine-learning
Storage details