Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Project: release a compact receipt model to an edge device

Last updated: 7 Oct 20265 min read
project
AdvancedBy AITrove Editorial

Package a receipt-quality model for a constrained device only after converted-runtime measurements and defect-specific decision gates pass.

Freeze the baseline contract

Pin the original receipt model, preprocessing revision, class map, confidence threshold and device runtime. Capture a held-out set grouped by physical receipt and a separate calibration population spanning lighting, blur, crop and device differences. Record model bytes, resident memory, cold-start time, median and 95th-percentile request latency at the intended batch size. Report class recall and false accepts. A smaller file by itself is not a release result. The classifier project supplies the decision context.

Build separate compression candidates

Create one structured-pruned candidate with a physically smaller graph and one calibrated quantized candidate in a backend that supports its operators. Optionally try a compact model with both changes, but recalibrate after pruning. Store each model and conversion configuration under distinct revisions. Do not reuse scales from the original graph after channels are removed. Structural parity checks the graph rewrite; calibration checks integer ranges.

Benchmark the actual serving path

Run the same normalized fixed inputs through baseline and candidates in the target runtime. Compare logits, decisions near the threshold and nonfinite outputs before bulk timing. Warm each candidate and measure many realistic requests, including preprocessing and serialization when they are part of user-visible latency. Record thermal state, power mode and backend version. An isolated desktop matrix benchmark cannot establish mobile or embedded tail latency. Report accepted-case coverage and manual review volume if the system can abstain.

Apply a predeclared release gate

Write gates before reading the final test results: for example, defect recall may not fall by more than one percentage point, false accepts may not rise, and the target device 95th-percentile latency must meet its limit. Compute uncertainty intervals or at least event counts for small slices; one error out of a tiny group is too noisy for a strong claim. The code below evaluates supplied metrics and returns all failed gates. It does not fabricate benchmark measurements. If no candidate passes, retain the baseline.

Ship with recovery state

The release bundle includes model, scales, preprocessing, class map, threshold, runtime compatibility and checksum. Run a clean-process reload test and a staged device rollout. Monitor out-of-range activation counts, latency tails, confidence coverage, manual review workload and class-specific corrections. Keep the previous package available for rollback if a source-device shift appears after deployment. The confidence audit identifies cases that should route to review rather than receive an automatic answer.

Implementation

python
from dataclasses import dataclass

@dataclass(frozen=True)
class EdgeMetrics:
    defect_recall: float
    false_accepts: int
    p95_latency_ms: float
    resident_megabytes: float

def release_failures(baseline: EdgeMetrics, candidate: EdgeMetrics,
                     latency_limit_ms: float, memory_limit_mb: float) -> list[str]:
    failures = []
    if candidate.defect_recall < baseline.defect_recall - 0.01:
        failures.append("defect recall gate")
    if candidate.false_accepts > baseline.false_accepts:
        failures.append("false accept gate")
    if candidate.p95_latency_ms > latency_limit_ms:
        failures.append("latency gate")
    if candidate.resident_megabytes > memory_limit_mb:
        failures.append("memory gate")
    return failures

reference = EdgeMetrics(0.94, 3, 31.4, 42.0)
converted = EdgeMetrics(0.925, 3, 18.7, 19.0)
assert release_failures(reference, converted, 23.0, 24.0) == [
    "defect recall gate"]

Performance and operating cost

Compression preparation can take longer than original training once calibration, fine-tuning and device validation are counted. Disk bytes are only one component of memory; activations, backend buffers and image preprocessing remain. For N measured requests, sorting latency for an exact empirical percentile costs O(N log N), while streaming approximations can lower memory. The release gate is O(1) once reliable metrics exist. Budget repeat device runs and human review of rare-class errors before claiming an edge win.

Common Mistakes

  • Do not publish a smaller checkpoint that runs floating-point fallback kernels and misses latency limits.
  • Do not set gates after inspecting the final test outcome.
  • Do not conceal a defect-recall failure with good average accuracy or a lower p95 latency.

Read next

Continue the workflow: Project: distill and release a receipt-quality student.

Continue the workflow: Project: stress-test receipt decisions under bounded image changes.

ai-data
deep-learning
Storage details