A cascade must meet branch-specific quality and a full-request deadline, including when a component fails.
Cascade release gates: branch quality, deadlines and cost
Budget every branch
Write a deadline for the complete receipt decision, then divide it among preprocessing, first-stage scoring, specialist scoring, routing overhead and fallback. Do not promise the specialist its full standalone timeout after the first stage has already consumed most of the request budget. Measure p99 end-to-end by branch, not an average of component p99 values. Tail observation shows the experienced deadline; degraded-path tests exercise failures.
Gate quality at route boundaries
Include receipts just above and below the first-stage cutoff, low-quality scans, large purchase amounts and missing fields. Recalculate which cases reach the specialist when either model or cutoff changes. If fewer cases reach it, verify whether missed-risk rate rises in the fast branch. A branch with little traffic may need a wider uncertainty interval and a longer canary. Slice gates protect small but important cohorts.
Count capacity and economic effects
Estimate cost per completed decision from first-stage work, specialist route share, retries and manual fallbacks. The specialist may need reserved capacity for a sudden shift in route share; autoscaling based only on total request rate can lag. Distinguish model budget from reviewer budget. The cost ledger provides an allocation, while multi-model admission protects shared host capacity.
Test promotion and rollback
Warm both components, validate interstage schema, inject a first-stage timeout and a specialist crash, and check the selected fallback. Stage the manifest as one versioned unit. Monitor branch share, missing stages, route errors, p99 and outcome coverage before expanding. A failed canary restores the prior manifest and its policy, not just the model that most recently changed. The project applies these gates to a receipt-quality plus risk cascade.
Implementation
def cascade_release_gate(quality_ok, branch_p99_ms, branch_budget_ms,
missed_fast_cases, max_missed_fast_cases):
if not quality_ok:
return "hold:quality"
if not branch_p99_ms or set(branch_p99_ms) != set(branch_budget_ms):
return "hold:branch-evidence"
if any(branch_p99_ms[name] > branch_budget_ms[name]
for name in branch_p99_ms):
return "hold:deadline"
if missed_fast_cases > max_missed_fast_cases:
return "hold:fast-branch-errors"
return "stage:cascade"
latency = {"fast": 41, "specialist": 86, "fallback": 70}
budget = {"fast": 57, "specialist": 94, "fallback": 82}
assert cascade_release_gate(True, latency, budget, 2, 3) == "stage:cascade"
assert cascade_release_gate(True, {**latency, "specialist": 101},
budget, 2, 3) == "hold:deadline"
Performance and operating cost
The branch gate is O(b) time and O(1) extra space for b branches. Benchmarks and adjudicated quality checks dominate operational cost. Specialist reserve capacity raises idle spend, but underprovisioning can turn a route-share shift into a deadline failure; size it from burst and branch-share observations.
Common Mistakes
- Adding component p99 values to claim an end-to-end p99.
- Checking only the common fast branch.
- Reducing specialist traffic without testing missed risk.
- Restoring an old model with a new routing cutoff after a failed canary.
Read next
- Model cascades: pin component versions and route decisions
- Project: release a receipt-quality and risk-model cascade
- Inference latency budgets: measure queue, feature and model time
- Slice quality gates when labels are sparse or delayed
- Inference cost ledger: allocate shared capacity without false precision
