Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Hosted-model migration: paired replay and reversible cutover

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Replay the same workload through incumbent and successor deployments, then move traffic only after behavior and capacity gates pass.

Build a paired evaluation frame

Take a consented, redacted set of multilingual support tickets with adjudicated destination queues. Replay each ticket through incumbent and candidate using pinned prompt, schema and decoding settings appropriate to each model. Store deployment identity, request ID, latency, output validity, route and abstention reason. Because outputs may vary, compare task decisions and contract compliance rather than exact strings. The dependency manifest identifies both deployments and their service limits.

Gate behavior by cohort

Compute route accuracy, unsafe misroutes, malformed JSON, abstention rate and p95 latency by language and ticket type. A candidate that improves overall accuracy but misroutes urgent Spanish safety tickets fails the protected slice. Count malformed responses as failures rather than dropping them from the denominator. Reconcile usage and billable units separately from model quality. Slice gates keep a minority failure visible.

Cut over in stages

Start with shadow requests where customer-facing routing still uses the incumbent. Then move a small supported cohort through a versioned pointer. Log provider deployment, prompt, parser and policy revisions on every ticket. Keep incumbent capacity and rollback path until the soak window closes. If the provider retires it before rollback is possible, preapprove a human-routing fallback; do not claim rollback remains available. Canary practice supplies staged exposure, and fallback policy bounds customer delay.

Retire the old path deliberately

After passing mature review outcomes, client contract tests and capacity checks, stop the old traffic, archive evaluation evidence, remove stale secrets and update consumer inventory. Do not delete incident evidence needed to explain tickets routed during the overlap. If an alias auto-upgrades, re-run probes and gate the resolved model revision before treating it as equivalent. The project rehearses a late retirement notice and an urgent-ticket regression.

Implementation

python
def hosted_candidate_gate(report, limits):
    if report["schema_failures"]:
        return "hold:response-contract"
    if report["urgent_slice_misses"] > limits["urgent_slice_misses"]:
        return "hold:urgent-slice"
    if report["p95_ms"] > limits["p95_ms"]:
        return "hold:latency"
    if report["peak_rps_capacity"] < limits["required_rps"]:
        return "hold:capacity"
    return "canary:versioned-pointer"

limits = {"urgent_slice_misses": 0, "p95_ms": 820,
          "required_rps": 47}
report = {"schema_failures": 0, "urgent_slice_misses": 1,
          "p95_ms": 680, "peak_rps_capacity": 63}
assert hosted_candidate_gate(report, limits) == "hold:urgent-slice"
assert hosted_candidate_gate({**report, "urgent_slice_misses": 0}, limits)        == "canary:versioned-pointer"

Performance and operating cost

The aggregate gate is O(1) time and space. Paired replay approximately doubles calls over the sampled frame, and keeping both deployments consumes overlapping capacity. Staged rollout increases operational coordination but lets teams stop a bad candidate before every ticket reaches it.

Common Mistakes

  • Dropping malformed candidate responses from quality denominators.
  • Comparing only average routing accuracy across languages.
  • Claiming rollback after the incumbent deployment has retired.
  • Switching an alias without recording the resolved deployment revision.

Read next

ai-data
mlops
Storage details