Replay the same workload through incumbent and successor deployments, then move traffic only after behavior and capacity gates pass.
Hosted-model migration: paired replay and reversible cutover
Build a paired evaluation frame
Take a consented, redacted set of multilingual support tickets with adjudicated destination queues. Replay each ticket through incumbent and candidate using pinned prompt, schema and decoding settings appropriate to each model. Store deployment identity, request ID, latency, output validity, route and abstention reason. Because outputs may vary, compare task decisions and contract compliance rather than exact strings. The dependency manifest identifies both deployments and their service limits.
Gate behavior by cohort
Compute route accuracy, unsafe misroutes, malformed JSON, abstention rate and p95 latency by language and ticket type. A candidate that improves overall accuracy but misroutes urgent Spanish safety tickets fails the protected slice. Count malformed responses as failures rather than dropping them from the denominator. Reconcile usage and billable units separately from model quality. Slice gates keep a minority failure visible.
Cut over in stages
Start with shadow requests where customer-facing routing still uses the incumbent. Then move a small supported cohort through a versioned pointer. Log provider deployment, prompt, parser and policy revisions on every ticket. Keep incumbent capacity and rollback path until the soak window closes. If the provider retires it before rollback is possible, preapprove a human-routing fallback; do not claim rollback remains available. Canary practice supplies staged exposure, and fallback policy bounds customer delay.
Retire the old path deliberately
After passing mature review outcomes, client contract tests and capacity checks, stop the old traffic, archive evaluation evidence, remove stale secrets and update consumer inventory. Do not delete incident evidence needed to explain tickets routed during the overlap. If an alias auto-upgrades, re-run probes and gate the resolved model revision before treating it as equivalent. The project rehearses a late retirement notice and an urgent-ticket regression.
Implementation
def hosted_candidate_gate(report, limits):
if report["schema_failures"]:
return "hold:response-contract"
if report["urgent_slice_misses"] > limits["urgent_slice_misses"]:
return "hold:urgent-slice"
if report["p95_ms"] > limits["p95_ms"]:
return "hold:latency"
if report["peak_rps_capacity"] < limits["required_rps"]:
return "hold:capacity"
return "canary:versioned-pointer"
limits = {"urgent_slice_misses": 0, "p95_ms": 820,
"required_rps": 47}
report = {"schema_failures": 0, "urgent_slice_misses": 1,
"p95_ms": 680, "peak_rps_capacity": 63}
assert hosted_candidate_gate(report, limits) == "hold:urgent-slice"
assert hosted_candidate_gate({**report, "urgent_slice_misses": 0}, limits) == "canary:versioned-pointer"
Performance and operating cost
The aggregate gate is O(1) time and space. Paired replay approximately doubles calls over the sampled frame, and keeping both deployments consumes overlapping capacity. Staged rollout increases operational coordination but lets teams stop a bad candidate before every ticket reaches it.
Common Mistakes
- Dropping malformed candidate responses from quality denominators.
- Comparing only average routing accuracy across languages.
- Claiming rollback after the incumbent deployment has retired.
- Switching an alias without recording the resolved deployment revision.
Read next
- Hosted-model dependencies: pin behavior and plan retirement
- Project: migrate a multilingual ticket router before model retirement
- Slice quality gates when labels are sparse or delayed
- Shadow and canary rollout: compare a candidate without losing a rollback
- Serving overload: bound queues and choose a fallback before time runs out
