Attach an adapter to a frozen event encoder, compare it with frozen and full-update baselines, then release with merge parity and base identity.
Project: adapt a service-event classifier with low-rank weights
Split whole incidents
Collect event text and structured fields with incident ID, source, event time and reviewed class. Hold out whole incidents and a later period; strip resolution notes and post-alert actions that reveal the answer. A small adapter can memorize leakage as easily as a full model. The split lesson protects the comparison across model choices.
Compare three candidates
Measure a frozen encoder with a trainable head, selected low-rank adapters, and fuller fine-tuning if resources allow. Use the same labels, tokenization, class weighting and held-out incidents. The snippet performs one synthetic update and confirms the base receives no weight gradient. It does not train a real event detector. The adapter lesson gives parameter and memory arithmetic.
Audit costly mistakes
Report severe-incident recall, macro F1, calibration and abstention by event family and source system. Compare class changes against the frozen baseline. A lower average loss can hide a drop on a rare incident class. Repeat across seeds, record trainable count and peak memory, and select the checkpoint with validation data rather than the final training step.
Check merge and serving
Pair the selected adapter with its exact base fingerprint. Compare unmerged and merged logits on fixed held-out events, then count thresholded decision flips. Test the deployed precision and runtime rather than extrapolating from float parity. Measure tokenization plus model p95 latency and resident memory for the adapter-loading strategy. The merge lesson prevents duplicate deltas and base drift.
Publish with fallback
Shadow-score incoming events and review disagreements before affecting any alert. Require no severe-recall drop, an acceptable review workload, matching base identity and declared merge tolerance. Retain the previous classifier if a source changes format or adapter load fails. Package base hash, target layers, rank, scale, tokenizer, threshold and rollback pointer as one release unit.
Implementation
import torch
from torch import nn
from torch.nn import functional as functional
torch.manual_seed(47)
base = nn.Linear(6, 8, bias=False)
base.weight.requires_grad_(False)
down, up = nn.Linear(6, 2, bias=False), nn.Linear(2, 8, bias=False)
nn.init.zeros_(up.weight)
head = nn.Linear(8, 3)
optimizer = torch.optim.AdamW(list(down.parameters()) +
list(up.parameters()) + list(head.parameters()), lr=.0007)
events = torch.rand(4, 6)
labels = torch.tensor([0, 2, 1, 2])
logits = head(torch.relu(base(events) + 2.5 * up(down(events))))
loss = functional.cross_entropy(logits, labels)
optimizer.zero_grad(set_to_none=True)
loss.backward()
optimizer.step()
assert torch.isfinite(loss) and base.weight.grad is None
assert up.weight.grad is not NonePerformance and operating cost
For an I-by-O frozen projection and rank R, the adapter adds R(I + O) trainable weights and O(BR(I + O)) work at batch B. The base still occupies IO weight storage and O(BIO) forward work, while larger encoders retain activation cost. The toy step checks gradient routing only; quality, merge parity and low-bit behavior require held-out incidents and target-runtime tests.
Common Mistakes
- Do not split messages from one incident between training and final test.
- Do not ship an adapter without exact base and tokenizer identity.
- Do not let overall accuracy hide severe-incident recall loss.
