Smaller integer weights are useful only when calibration data, operator support and held-out decisions survive conversion to the target runtime.
Post-training quantization and calibration contracts
Separate weight rounding from executable quantization
Quantization maps a real range to integer codes through a scale and sometimes a zero point. A symmetric signed eight-bit weight mapping can use a separate scale for each output channel. That reduces rounding error when channels have different ranges. The code below performs a transparent weight round-trip and measures reconstruction error. It does not produce an integer inference kernel: multiplying restored floating-point weights still runs floating-point arithmetic. A deployable conversion also needs supported operators, activation handling and a target-specific runtime. Inference contracts include these choices.
Choose representative calibration traffic
For post-training static activation quantization, observers estimate activation ranges from a representative calibration set. Select data after the same preprocessing used at deployment, covering normal receipts, blur, cut-off edges, exposure extremes and capture devices. Keep calibration separate from final test identities and labels used for release decisions. A set of random tensors can exercise code paths but cannot characterize real activation ranges. Record sample count, collection window, device mix, preprocessing hash and observer configuration. Drift in those inputs can make a previously acceptable scale harmful.
Inspect outliers and granularity
One extreme weight or activation can expand a tensor-wide scale so much that small meaningful values collapse to zero. Per-channel weight scales often improve resolution, but not every runtime accelerates every granularity. Compare per-tensor and per-channel error by layer and by input slice. Clipping a range can lower average error while saturating rare but important defects. Plot clipping frequency and held-out decision changes, especially for the examples closest to the review threshold. A confidence policy may be more sensitive than top-class accuracy.
Evaluate the converted artifact
Run the exported model in the actual device runtime, then compare outputs with the original model on fixed and held-out batches. Audit preprocessing parity, unsupported operator fallbacks, class-specific error, calibration and latency at realistic batch size. Count artifact bytes including scales and metadata, and measure resident memory rather than multiplying parameter count by a theoretical compression ratio. A model that silently dequantizes between layers may shrink on disk but miss the latency target. Keep baseline and candidate artifacts under versioned manifests for rollback.
Escalate only after diagnosis
If conversion damages a critical class, isolate the first layer with a large activation or output discrepancy. Try better calibration coverage, a different supported granularity or leaving that layer in higher precision. Quantization-aware training is a separate intervention that simulates rounding effects while updating weights; it has training cost and still needs held-out validation. Do not tune repeatedly on the final test. The edge project turns measurements into a release decision.
Implementation
import torch
classifier_weights = torch.tensor([[0.08, -0.32, 0.71, -0.19],
[2.4, -1.8, 0.3, 0.06],
[-0.04, 0.11, -0.09, 0.24]])
channel_maximum = classifier_weights.abs().amax(dim=1, keepdim=True)
channel_scales = channel_maximum.clamp_min(1e-8) / 127
integer_codes = (classifier_weights / channel_scales).round().clamp(
-127, 127).to(torch.int8)
restored_weights = integer_codes.float() * channel_scales
maximum_error = (classifier_weights - restored_weights).abs().amax(dim=1)
assert integer_codes.dtype == torch.int8
assert torch.isfinite(maximum_error).all()
assert bool((maximum_error <= channel_scales.squeeze(1) / 2 + 1e-6).all())Performance and operating cost
Computing per-channel maximums and rounding is O(P) in the number of weights. Raw signed eight-bit codes occupy roughly one byte per weight versus four bytes for float32, before channel scales, packing alignment and runtime metadata. Static activation calibration adds a forward pass over representative inputs; deployment conversion and benchmark cost depend on the target backend. End-to-end latency may improve, stay flat or regress because of unsupported operations and quantize/dequantize boundaries. Measure on the actual device.
Common Mistakes
- Do not claim integer inference from a dequantized-weight demonstration.
- Do not calibrate activations with random tensors or final-test identities.
- Do not hide rare-class regressions behind aggregate accuracy and file size.
Read next
- Inference contracts: preserve preprocessing and measure tail latency
- Structured pruning and latency parity
- Project: release a compact receipt model to an edge device
- Temperature scaling and selective risk
- Staged unfreezing and domain-shift audits
Continue the workflow: Project: distill and release a receipt-quality student.
Continue the workflow: LoRA merge parity, base identity and adapter release.
