A compression target should be defined by end-to-end latency, peak memory, artifact size and error on the deployment device before any model is changed.
Model inference budget and device profile
Measure the path users wait for
A depot camera model does not serve only a matrix multiplication. Image decode, crop, normalization, encoder, head, serialization and queue time all affect the moment an inspection result arrives. Set a decision deadline and profile the full pipeline on the intended device under realistic concurrency. The serving guide defines the input and timing boundary.
Report a distribution, not a fastest run
Warm up the runtime, then retain enough timed cases across image sizes and camera types. The code computes a nearest-rank median and 95th percentile from an illustrative set of completed requests; it is a statistics check, not an actual benchmark. Report sample count, hardware, batch size, thread count and cold-start behavior with the numbers. A compressed model that helps the median but misses the deadline for the largest images may not help operations.
Measure memory and storage separately
Artifact bytes affect delivery and cold start; peak resident memory includes activations, runtime and preprocessing buffers. Parameter count is neither peak memory nor end-to-end latency. A sparse weight tensor may serialize smaller while running at the same speed on a backend without sparse kernels. Pruning makes that distinction concrete.
Keep accuracy at the same decision point
The baseline and candidate need the same target cohort, preprocessing contract, threshold and action cost unless a policy change is explicitly part of the comparison. Report false negatives by depot and camera, not only average accuracy. Target-slice auditing can expose a small but costly regression.
Turn the budget into a test
Name maximum artifact size, peak memory, p95 latency and acceptable error before tuning compression settings. Use development data for choosing settings and a later untouched target period for acceptance. The Pareto review compares candidates; the project records promotion rules.
Implementation
from math import ceil
# Completed end-to-end inspection requests on one declared device profile.
latency_ms = [32, 35, 36, 37, 39, 40, 42, 43, 44, 45,
46, 47, 48, 49, 50, 52, 55, 58, 67, 74]
def nearest_rank(values, fraction):
if not values or not 0 < fraction <= 1:
raise ValueError("valid samples and fraction required")
ordered = sorted(values)
return ordered[ceil(fraction * len(ordered)) - 1]
profile = {"requests": len(latency_ms),
"median_ms": nearest_rank(latency_ms, 0.50),
"p95_ms": nearest_rank(latency_ms, 0.95),
"deadline_misses": sum(duration > 60 for duration in latency_ms)}
assert profile == {"requests": 20, "median_ms": 45,
"p95_ms": 67, "deadline_misses": 2}Performance and operating cost
Sorting N latency samples takes O(N log N) time and O(N) storage; streaming approximate quantiles can reduce memory for large telemetry. A valid device benchmark consumes repeated real inference and must include realistic concurrency and input sizes. Training or compressing a model can cost far more than this reporting step.
Common Mistakes
- Do not compare candidates timed on different hardware or input mixes.
- Do not call parameter count a measured memory or latency saving.
- Do not hide cold start or tail latency behind an average.
Read next
- Teacher–student distillation objective
- Structured pruning and compute shape
- Affine quantization and range audit
- Compression Pareto review and shadow check
Continue the workflow: Constrained sequence decoding and valid label paths.
