Skip to content
AITroveRead. Build. Understand.
Make this comfortable

LoRA merge parity, base identity and adapter release

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A merged adapter must reproduce the unmerged evaluation path within tolerance; base drift and duplicate application are release defects.

Merge with the declared scale

For base weight W, down matrix A, up matrix B and scale alpha divided by rank, the merged weight is W plus scale times B multiplied by A. In evaluation mode with dropout disabled, its linear output should match the separate adapter branch within numeric tolerance. The code compares both paths on fixed inputs. The rank lesson explains the factor dimensions.

Bind to the correct checkpoint

An adapter file is incomplete without the base model hash, target-layer names, tokenizer or feature schema and preprocessing revision. A new base with the same shapes can still invalidate the adapter. Save merge script revision as well. Checkpoint identity matters for a release as much as for training recovery. Refuse a load when the manifest does not match.

Do not apply twice

Adding an adapter delta to weights that were already merged doubles the update. Store an explicit merged or unmerged state and test the loader for repeat calls. Multiple task adapters need immutable base weights or isolated merged copies, not mutation of one shared process between requests. If requests can use different adapters concurrently, test that their outputs never cross.

Check target numeric precision

Merging before quantization need not match merging into an already quantized weight. Rounding, scale and adapter dropout can shift scores. Evaluate in deterministic mode at the deployed dtype, report maximum logit difference and count any class-decision flips. The calibration lesson supplies an additional quality check for low-bit serving.

Gate a reversible release

Measure task quality, p95 latency and peak memory for unmerged and merged paths. A merged copy removes the extra adapter branch at inference but may require one full model copy per task. Keep the previous base plus adapter as a rollback pair. The applied project compares both release paths on held-out incidents.

Implementation

python
import torch
from torch import nn

torch.manual_seed(47)
base = nn.Linear(6, 4, bias=False)
down = nn.Linear(6, 2, bias=False)
up = nn.Linear(2, 4, bias=False)
features = torch.rand(3, 6)
scale = 5 / 2
unmerged = base(features) + scale * up(down(features))
combined_weight = base.weight + scale * (up.weight @ down.weight)
merged = nn.functional.linear(features, combined_weight)
assert combined_weight.shape == (4, 6)
assert torch.allclose(unmerged, merged, atol=1e-6, rtol=1e-5)
assert not torch.allclose(base.weight, combined_weight)

Performance and operating cost

One I-by-O merge at rank R takes O(IOR) arithmetic and materializes O(IO) combined weights. Leaving the adapter separate adds O(BR(I + O)) work per batch B. Keeping several merged copies can multiply resident base memory by task count. Quantized merge arithmetic needs its own parity measurement; this example checks floating-point algebra only.

Common Mistakes

  • Do not merge an adapter twice.
  • Do not use a same-shaped but different base model.
  • Do not infer low-bit parity from a full-precision unit check.

Read next

ai-data
deep-learning
Storage details