Convert a receipt-risk classifier, test score and route parity, qualify hardware classes and stage a reversible rollout.
Project: release a faster receipt model without changing review routes
Prepare the release packet
Start with a pinned full-precision receipt-risk model, preprocessing digest and threshold policy. Create a candidate from a recorded conversion configuration and a calibration snapshot that excludes final evaluation receipts. Record the candidate digest and runtime build. The business contract is unchanged: receipts above the review threshold still go to review, and missing fields still take the existing fallback route. Numerical parity is a separate acceptance gate from hardware qualification.
Run the frozen replay
Replay 4,700 paired decisions covering ordinary scans, blurred images, short receipts, high-value purchases and a dense band around the review threshold. Preserve each decision ID so discrepancies are traceable. Compare maximum and percentile score error, route flips and slice-level quality. A candidate that reduces median inference from 51 to 34 milliseconds but moves four borderline receipts away from manual review is held; latency cannot waive the route contract. Investigate calibration coverage before changing the business threshold.
Qualify each deployment cell
Benchmark the corrected candidate on the production ARM CPU, a newer accelerator node and an older fallback node at fixed concurrency. Record selected provider, warm and cold latency, p99, peak memory and the baseline comparison. The older node dequantizes part of the graph and exceeds its p99 budget, so it keeps the full-precision artifact. Verify the baseline still loads after the runtime upgrade. Offline telemetry is relevant if the ARM cohort includes handhelds.
Ship with an evidence-backed pointer
Stage a small compatible cohort, compare observed route rates and latency by hardware class, then expand only after enough exposed decisions arrive. The release packet contains the conversion recipe, calibration and evaluation digests, hardware matrix, threshold-flip report, rollback digest and cohort observation window. Reversible promotion moves each class independently. A runtime or driver change invalidates the corresponding performance cell even when the artifact bytes do not change.
Implementation
def receipt_optimized_release(cells, route_flips):
if route_flips < 0:
raise ValueError("negative route flips")
if route_flips:
return "hold:route-parity"
if not cells:
return "hold:no-hardware-evidence"
if any(not cell["provider_verified"] or cell["p99_ms"] > cell["budget_ms"]
for cell in cells):
return "partial:baseline-on-failing-cells"
return "stage:candidate"
cells = [{"provider_verified": True, "p99_ms": 69, "budget_ms": 84},
{"provider_verified": False, "p99_ms": 91, "budget_ms": 84}]
assert receipt_optimized_release(cells, 4) == "hold:route-parity"
assert receipt_optimized_release(cells, 0) == "partial:baseline-on-failing-cells"
assert receipt_optimized_release(cells[:1], 0) == "stage:candidate"
Performance and operating cost
The disposition loop is O(h) time and O(1) extra space for h hardware cells; paired score replay is O(n) across n receipts. The large cost is executing both models, keeping representative snapshots and benchmarking every supported cell. Partial rollout preserves a storage copy of the baseline and adds a temporary mixed-fleet burden, but avoids treating an incompatible host as a release success.
Common Mistakes
- Waiving route flips because average accuracy stays flat.
- Claiming a fleet-wide speedup from one fast host.
- Changing the threshold to compensate for conversion drift.
- Omitting the baseline load test after a runtime upgrade.
Read next
- Optimized model release: calibration, score drift and parity
- Hardware matrix for optimized models: provider support and fallback
- Promotion evidence: bind evaluation, contract and rollback to one digest
- Slice quality gates when labels are sparse or delayed
- Offline edge telemetry: late events, coverage and rollback
