Package a receipt-quality model for a constrained device only after converted-runtime measurements and defect-specific decision gates pass.
Project: release a compact receipt model to an edge device
Freeze the baseline contract
Pin the original receipt model, preprocessing revision, class map, confidence threshold and device runtime. Capture a held-out set grouped by physical receipt and a separate calibration population spanning lighting, blur, crop and device differences. Record model bytes, resident memory, cold-start time, median and 95th-percentile request latency at the intended batch size. Report class recall and false accepts. A smaller file by itself is not a release result. The classifier project supplies the decision context.
Build separate compression candidates
Create one structured-pruned candidate with a physically smaller graph and one calibrated quantized candidate in a backend that supports its operators. Optionally try a compact model with both changes, but recalibrate after pruning. Store each model and conversion configuration under distinct revisions. Do not reuse scales from the original graph after channels are removed. Structural parity checks the graph rewrite; calibration checks integer ranges.
Benchmark the actual serving path
Run the same normalized fixed inputs through baseline and candidates in the target runtime. Compare logits, decisions near the threshold and nonfinite outputs before bulk timing. Warm each candidate and measure many realistic requests, including preprocessing and serialization when they are part of user-visible latency. Record thermal state, power mode and backend version. An isolated desktop matrix benchmark cannot establish mobile or embedded tail latency. Report accepted-case coverage and manual review volume if the system can abstain.
Apply a predeclared release gate
Write gates before reading the final test results: for example, defect recall may not fall by more than one percentage point, false accepts may not rise, and the target device 95th-percentile latency must meet its limit. Compute uncertainty intervals or at least event counts for small slices; one error out of a tiny group is too noisy for a strong claim. The code below evaluates supplied metrics and returns all failed gates. It does not fabricate benchmark measurements. If no candidate passes, retain the baseline.
Ship with recovery state
The release bundle includes model, scales, preprocessing, class map, threshold, runtime compatibility and checksum. Run a clean-process reload test and a staged device rollout. Monitor out-of-range activation counts, latency tails, confidence coverage, manual review workload and class-specific corrections. Keep the previous package available for rollback if a source-device shift appears after deployment. The confidence audit identifies cases that should route to review rather than receive an automatic answer.
Implementation
from dataclasses import dataclass
@dataclass(frozen=True)
class EdgeMetrics:
defect_recall: float
false_accepts: int
p95_latency_ms: float
resident_megabytes: float
def release_failures(baseline: EdgeMetrics, candidate: EdgeMetrics,
latency_limit_ms: float, memory_limit_mb: float) -> list[str]:
failures = []
if candidate.defect_recall < baseline.defect_recall - 0.01:
failures.append("defect recall gate")
if candidate.false_accepts > baseline.false_accepts:
failures.append("false accept gate")
if candidate.p95_latency_ms > latency_limit_ms:
failures.append("latency gate")
if candidate.resident_megabytes > memory_limit_mb:
failures.append("memory gate")
return failures
reference = EdgeMetrics(0.94, 3, 31.4, 42.0)
converted = EdgeMetrics(0.925, 3, 18.7, 19.0)
assert release_failures(reference, converted, 23.0, 24.0) == [
"defect recall gate"]Performance and operating cost
Compression preparation can take longer than original training once calibration, fine-tuning and device validation are counted. Disk bytes are only one component of memory; activations, backend buffers and image preprocessing remain. For N measured requests, sorting latency for an exact empirical percentile costs O(N log N), while streaming approximations can lower memory. The release gate is O(1) once reliable metrics exist. Budget repeat device runs and human review of rare-class errors before claiming an edge win.
Common Mistakes
- Do not publish a smaller checkpoint that runs floating-point fallback kernels and misses latency limits.
- Do not set gates after inspecting the final test outcome.
- Do not conceal a defect-recall failure with good average accuracy or a lower p95 latency.
Read next
- Post-training quantization and calibration contracts
- Structured pruning and latency parity
- Inference contracts: preserve preprocessing and measure tail latency
- Project: audit receipt confidence and review handoff
- Project: classify receipt image quality with a checked training contract
Continue the workflow: Project: distill and release a receipt-quality student.
Continue the workflow: Project: stress-test receipt decisions under bounded image changes.
