Quantization maps floating values to a finite integer range using a scale and zero point, trading representation precision for possible model-size and device-speed gains.
Affine quantization and range audit
Understand the mapping
For a chosen floating range, a uniform integer quantizer rounds each value divided by a scale and shifted by a zero point, then clamps to supported integer limits. Dequantization reverses the mapping approximately. The code uses a symmetric signed eight-bit teaching variant with zero point zero. A production converter chooses ranges and operator behavior for its backend; copying this small function is not an export pipeline. The device profile tells whether a lower-precision artifact helps.
Choose calibration data with coverage
Post-training activation quantization needs a representative set of inputs to observe value ranges. If calibration sees only bright parcel photos, darker depot images may clip. Measure saturation and output drift by camera and package type. Never select the calibration range on the final target test. Target grouping preserves the later check.
Separate size from accuracy and speed
Eight-bit weights can use fewer bytes than 32-bit floats, but extra scales, unsupported operators or conversion steps affect both size and latency. Some hardware accelerates the integer path; other devices may not. Report actual artifact bytes, p95 latency and peak memory on the serving backend, with the same preprocessing as the baseline. The Pareto review keeps these measurements together.
Investigate outliers and rare errors
A wide calibration range may waste precision near common values; a tight one clips rare extremes. Quantization can move a parcel score across an action threshold even when average numerical error is small. Compare issued decisions and per-label false negatives, not only mean absolute output difference. Action costs identify consequential changes.
Escalate only after measured failure
If post-training conversion causes unacceptable target error, consider another calibration policy or quantization-aware training, selected on development data. This increases training and maintenance cost; it is not automatically better. The release review records the exact converter, range and rollback artifact.
Implementation
from math import isfinite
calibration_peak = 2.54
scale = calibration_peak / 127
def symmetric_int8(values, step):
if step <= 0:
raise ValueError("positive scale required")
encoded = []
for value in values:
if not isfinite(value):
raise ValueError("finite input required")
encoded.append(max(-127, min(127, round(value / step))))
return encoded, [integer * step for integer in encoded]
inspection_values = [0.0, 0.37, -1.14, 2.8]
integers, restored = symmetric_int8(inspection_values, scale)
clipped = sum(abs(value / scale) > 127 for value in inspection_values)
assert integers[0] == 0
assert integers[-1] == 127
assert clipped == 1
assert abs(restored[1] - inspection_values[1]) <= scale / 2 + 1e-12Performance and operating cost
Converting N scalar values costs O(N) time and O(N) output storage; packed integer model inference depends on backend support. Calibration scans representative data, and quantization-aware training adds model optimization work. Measure transfer and conversion overhead, because integer arithmetic alone does not determine end-to-end latency.
Common Mistakes
- Do not calibrate activation ranges on the sealed final test.
- Do not equate fewer weight bits with a measured latency gain.
- Do not ignore threshold flips on rare costly cases.
