A converted model can run slower, use more memory or fail on a different processor; qualify the actual serving fleet.
Hardware matrix for optimized models: provider support and fallback
Name the real target fleet
Inventory CPU instruction class, accelerator, runtime build, execution provider, memory limit and app version for each serving cohort. “GPU” is not a hardware specification. Some hosts can execute a quantized graph but lack the fast operator implementation; they silently dequantize or move part of the graph to the CPU. Record the actual provider selected at load time, not merely the preferred provider in configuration. Edge release manifests need the same capability check.
Benchmark the full request path
Measure warm and cold initialization, preprocessing, inference, postprocessing, peak memory, p50 and tail latency at a defined concurrency. Compare with the existing artifact on identical input batches and host classes. Include throttled devices and small batches; a throughput win on a full accelerator can become a tail-latency loss at low traffic. Track fallback count and provider partition changes. Inference chain identity makes the two benchmarks comparable.
Route unsupported hosts explicitly
The release manifest maps each qualified hardware class to a tested artifact and runtime. Unknown classes retain the approved baseline or enter a controlled hold. Do not assume an optimized artifact is universally smaller in memory after runtime graph expansion. A class that fails correctness or latency must not receive the candidate, even if the aggregate fleet metric passes. Preserve a rollback package on each host and test that it loads under the current runtime. The promotion pointer must resolve to a compatible artifact.
Watch the deployed execution path
Emit model digest, runtime revision, selected provider and hardware class with sampled decision telemetry. A new host image or driver can change provider choice without changing the model digest. Alert on unexpected fallback and tail-latency regressions by class; a single global average conceals both. The project qualifies three host classes and deliberately holds one that loses its accelerator path. Re-run the matrix when runtime, driver or graph settings change.
Implementation
def qualified_artifact(host, matrix, baseline_digest):
key = (host["hardware"], host["runtime"], host["provider"])
qualification = matrix.get(key)
if qualification is None or not qualification["parity_passed"]:
return baseline_digest
if qualification["p99_ms"] > qualification["p99_budget_ms"]:
return baseline_digest
return qualification["artifact_digest"]
matrix = {("arm-r5", "ort-r8", "cpu-int8"):
{"parity_passed": True, "p99_ms": 63, "p99_budget_ms": 82,
"artifact_digest": "receipt-int8-47"}}
host = {"hardware": "arm-r5", "runtime": "ort-r8", "provider": "cpu-int8"}
assert qualified_artifact(host, matrix, "receipt-fp32-31") == "receipt-int8-47"
assert qualified_artifact({**host, "provider": "cpu-fallback"}, matrix,
"receipt-fp32-31") == "receipt-fp32-31"
Performance and operating cost
Lookup is O(1) expected time and space after a qualification matrix exists. Matrix construction scales with artifact variants, hardware classes and workloads; each cell needs correctness and performance runs. Retaining a baseline costs storage but bounds an unsupported-provider failure. More variants can improve fleet coverage while increasing validation and rollback work.
Common Mistakes
- Reading configured provider as proof the provider actually executed.
- Using one accelerator benchmark for all host classes.
- Ignoring model-load memory and cold-start time.
- Letting unknown hardware receive an unqualified optimized artifact.
Read next
- Optimized model release: calibration, score drift and parity
- Project: release a faster receipt model without changing review routes
- Edge model releases: pin runtime, preprocessing and cohort
- Multi-stage inference: pin each stage and its contract
- Promotion evidence: bind evaluation, contract and rollback to one digest
Continue the workflow: Project: patch a receipt-model loader across mixed hosts.
