Reconcile billing and completed decisions, identify retry amplification and choose a budget response that preserves quality.
Project: explain a receipt-model serving cost spike
Reconstruct the window
A receipt-risk endpoint appears to cost 38 percent more after a release. Freeze the before and after windows, model and image digests, completed-decision counts, instance prices, retry policy and traffic mix. Billing shows more replica-hours; traces show a higher request count but not a matching increase in completed receipts. Build the ledger from resource cost and completion identity before saying the model became more expensive. Allocation rules must leave idle and unknown capacity visible.
Find the amplification
Group attempts by idempotency key and record retries, timeouts and successful completions. One client cohort retries after 47 milliseconds, earlier than the service’s normal tail, so many receipts are scored twice. Separate retry work from model compute and compare p99 by host class. The new model also falls back on an older CPU class; that is a second, narrower cost driver. The result ledger prevents duplicate attempts from being counted as new business decisions.
Test safe responses
Correct the client retry deadline and verify idempotent completion. Keep the old host class on its qualified artifact until the optimized path passes its hardware matrix. Recalculate forecast cost per completed decision under the expected mix, then check quality and p99 against the baseline. Do not lower the review threshold or drop blurry scans just to meet the budget. The budget gate requires all dimensions to pass.
Close the accounting loop
Publish observed resource cost, completed decisions, duplicate attempts, provider fallback, unallocated spend, quality and p99. Assign separate owners to retry logic and runtime qualification. Stage each correction so the cost effect can be measured without confusing it with traffic changes. Keep an expiry on any budget exception and a rollback pointer for the serving artifact. Promotion identity makes the next comparison auditable.
Implementation
def completed_unit_cost(resource_cost, attempts):
if resource_cost < 0:
raise ValueError("negative cost")
completed = {attempt["decision_id"] for attempt in attempts
if attempt["state"] == "completed"}
if not completed:
return {"cost_each": None, "completed": 0}
return {"cost_each": resource_cost / len(completed),
"completed": len(completed)}
attempts = [{"decision_id": "receipt-47", "state": "timeout"},
{"decision_id": "receipt-47", "state": "completed"},
{"decision_id": "receipt-82", "state": "completed"}]
assert completed_unit_cost(0.14, attempts)["completed"] == 2
assert completed_unit_cost(0.14, attempts)["cost_each"] == 0.07
assert completed_unit_cost(0.14, attempts[:1])["cost_each"] is None
Performance and operating cost
Deduplicating a attempts uses O(a) expected time and O(d) memory for d distinct completed decisions. At production scale, do this in a bounded window or keyed store rather than loading every event into process memory. A corrected retry policy saves wasted compute but must preserve completion reliability; test timeout and replay semantics before rollout.
Common Mistakes
- Dividing the bill by all attempts instead of unique completions.
- Blaming the model before isolating client retries and provider fallback.
- Ignoring idle replica-hours in the cost change.
- Meeting the budget by dropping difficult cases or changing safety routes.
Read next
- Inference cost ledger: allocate shared capacity without false precision
- Inference budget gates: cost, latency and quality together
- Async result ledgers: reconcile output, failure and notification
- Hardware matrix for optimized models: provider support and fallback
- Promotion evidence: bind evaluation, contract and rollback to one digest
