Skip to content
AITroveRead. Build. Understand.
Make this comfortable

LoRA rank, scaling and frozen-base contracts

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Low-rank adapters train two small matrices around a frozen map; their savings depend on target layers, activations and exact base identity.

Count the changed weights

For a projection from 47 inputs to 83 outputs, a rank-four adapter trains four times the sum of 47 and 83, or 520 matrix entries. The base map holds 47 times 83, or 3,901 entries, and stays resident even when frozen. The effective update is a scaled product of the two adapter factors. Record any trainable biases or task head separately; freezing one layer is not freezing the entire encoder.

Choose target projections

Adapters can modify attention query and value projections, output projections or feed-forward maps. More targets increase trainable memory and capacity. Start with a measured subset, then compare held-out quality and peak memory against full fine-tuning and a frozen-feature baseline. The transfer lesson frames that choice. A small rank can underfit a task; it is not a universal defense against overfitting.

Initialize without a surprise change

A zero-initialized output factor makes the adapted model initially match the frozen base. Scaling, often alpha divided by rank, sets the update magnitude and belongs in saved metadata. With both factors initialized nonzero, the first forward already differs from the base. The snippet verifies base-output parity at initialization and confirms only adapter weights receive gradients. Compare several seeds when selecting rank and scaling.

Profile actual memory

Freezing base weights avoids their gradients and optimizer-state tensors, yet the base weights and forward activations still consume memory. Long sequences can make activations the dominant expense. Quantizing the frozen base changes arithmetic and needs separate calibration; it is not merely a smaller file. The memory lesson separates reserved memory from active tensors.

Keep the adapter attached to its base

Save the exact base-checkpoint fingerprint, targeted layer names, rank, scaling, dropout, tokenizer and feature schema with adapter weights. A different base can accept the same tensor shapes while producing incorrect outputs. The merge lesson checks that the exported combined weight reproduces the unmerged evaluation path.

Implementation

python
import torch
from torch import nn

torch.manual_seed(47)
base = nn.Linear(47, 83, bias=False)
base.weight.requires_grad_(False)
down = nn.Linear(47, 4, bias=False)
up = nn.Linear(4, 83, bias=False)
nn.init.zeros_(up.weight)
features = torch.rand(2, 47)
adapted = base(features) + (8 / 4) * up(down(features))
assert torch.allclose(adapted, base(features))
adapted.square().mean().backward()
assert base.weight.grad is None and up.weight.grad is not None
assert 4 * (47 + 83) == 520 and 47 * 83 == 3901

Performance and operating cost

A dense projection costs O(BIO) operations for batch B, input I and output O. The adapter adds O(BR(I + O)) work and R(I + O) trainable weights for rank R. The base still consumes IO stored weights and its forward compute; optimizer memory falls with trainable parameters, not total parameters. Whole-model savings require profiling target layers and sequence lengths.

Common Mistakes

  • Do not claim frozen base weights disappear from memory.
  • Do not omit target names, rank or scale from an adapter manifest.
  • Do not assume every task can match full fine-tuning with a tiny rank.

Read next

ai-data
deep-learning
Storage details