Review a smaller parcel-inspection model against the same target cases and serving device, with accuracy, memory, latency, shadow behavior and rollback.
Model compression release project
Set the deployment budget
Record the camera input contract, target label definitions, decision deadline, p95 latency limit, peak memory cap and artifact size limit. Profile the complete existing service before choosing a compression method. The baseline guide defines the measurements.
Produce candidate artifacts
A student may train from teacher outputs; a structured-pruned model removes connected units; a quantized model changes numeric representation. Each method needs versioned settings and a target development cohort. The methods can be combined, but every combination is a new candidate requiring its own validation. Distillation, pruning and quantization have different failure modes.
Compare outcomes and cost
Use a sealed target period and the same device to report per-label misses, calibration, p95 latency, peak memory, artifact bytes and deadline failures. Evaluate site and image-size slices. A nondominated candidate is not automatically acceptable if it breaches a hard miss or capacity bound. The Pareto review makes the tradeoffs visible.
Shadow before a bounded pilot
Log candidate and live predictions on the same requests without changing actions. Count missing candidate scores and runtime errors. Inspect later mature outcomes, then plan any action-changing pilot with an owner and retained old artifact. Shadow comparison is evidence about the current policy, not a full counterfactual policy test.
Record what remains unresolved
The code checks that a release packet includes required artifacts. It cannot certify the model or make an acceptance decision from absent numbers. Attach measured results and limits. If device testing or rare-label review is incomplete, hold promotion.
Implementation
def compression_release_gate(packet):
required = {
"device_profile_fixed": "device profile",
"target_test_sealed": "target test",
"rare_label_errors_checked": "rare-label errors",
"p95_and_memory_measured": "latency and memory",
"artifact_compatibility_checked": "artifact compatibility",
"shadow_failures_reported": "shadow failures",
"rollback_artifact_retained": "rollback artifact",
}
missing = [label for field, label in required.items() if not packet.get(field)]
return "eligible for pilot review" if not missing else "hold: " + ", ".join(missing)
inspection_packet = {
"device_profile_fixed": True, "target_test_sealed": True,
"rare_label_errors_checked": False,
"p95_and_memory_measured": True,
"artifact_compatibility_checked": True,
"shadow_failures_reported": False,
"rollback_artifact_retained": True,
}
assert compression_release_gate(inspection_packet) == (
"hold: rare-label errors, shadow failures"
)Performance and operating cost
The gate is O(K) for K checks. Distillation adds teacher inference and student training; pruning and quantization add conversion and validation. End-to-end benchmarks and a shadow period use real serving resources, and enough target labels are needed to assess rare costly regressions.
Common Mistakes
- Do not claim a model is faster from size or parameter count alone.
- Do not accept pooled accuracy when rare costly labels regress.
- Do not discard the previous serving artifact before rollback is tested.
Read next
- Model inference budget and device profile
- Teacher–student distillation objective
- Structured pruning and compute shape
- Affine quantization and range audit
- Compression Pareto review and shadow check
- Multi-label parcel review project
Continue the workflow: Project: release review for shared pump-inspection heads.
