A merged adapter must reproduce the unmerged evaluation path within tolerance; base drift and duplicate application are release defects.
LoRA merge parity, base identity and adapter release
Merge with the declared scale
For base weight W, down matrix A, up matrix B and scale alpha divided by rank, the merged weight is W plus scale times B multiplied by A. In evaluation mode with dropout disabled, its linear output should match the separate adapter branch within numeric tolerance. The code compares both paths on fixed inputs. The rank lesson explains the factor dimensions.
Bind to the correct checkpoint
An adapter file is incomplete without the base model hash, target-layer names, tokenizer or feature schema and preprocessing revision. A new base with the same shapes can still invalidate the adapter. Save merge script revision as well. Checkpoint identity matters for a release as much as for training recovery. Refuse a load when the manifest does not match.
Do not apply twice
Adding an adapter delta to weights that were already merged doubles the update. Store an explicit merged or unmerged state and test the loader for repeat calls. Multiple task adapters need immutable base weights or isolated merged copies, not mutation of one shared process between requests. If requests can use different adapters concurrently, test that their outputs never cross.
Check target numeric precision
Merging before quantization need not match merging into an already quantized weight. Rounding, scale and adapter dropout can shift scores. Evaluate in deterministic mode at the deployed dtype, report maximum logit difference and count any class-decision flips. The calibration lesson supplies an additional quality check for low-bit serving.
Gate a reversible release
Measure task quality, p95 latency and peak memory for unmerged and merged paths. A merged copy removes the extra adapter branch at inference but may require one full model copy per task. Keep the previous base plus adapter as a rollback pair. The applied project compares both release paths on held-out incidents.
Implementation
import torch
from torch import nn
torch.manual_seed(47)
base = nn.Linear(6, 4, bias=False)
down = nn.Linear(6, 2, bias=False)
up = nn.Linear(2, 4, bias=False)
features = torch.rand(3, 6)
scale = 5 / 2
unmerged = base(features) + scale * up(down(features))
combined_weight = base.weight + scale * (up.weight @ down.weight)
merged = nn.functional.linear(features, combined_weight)
assert combined_weight.shape == (4, 6)
assert torch.allclose(unmerged, merged, atol=1e-6, rtol=1e-5)
assert not torch.allclose(base.weight, combined_weight)Performance and operating cost
One I-by-O merge at rank R takes O(IOR) arithmetic and materializes O(IO) combined weights. Leaving the adapter separate adds O(BR(I + O)) work per batch B. Keeping several merged copies can multiply resident base memory by task count. Quantized merge arithmetic needs its own parity measurement; this example checks floating-point algebra only.
Common Mistakes
- Do not merge an adapter twice.
- Do not use a same-shaped but different base model.
- Do not infer low-bit parity from a full-precision unit check.
