Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Shadow and canary rollout: compare a candidate without losing a rollback

Last updated: 7 Oct 20265 min read
tutorial
IntermediateBy AITrove Editorial

A release stages candidate observation, limited live traffic and rollback against explicit health and decision metrics.

Use shadow for observation

Send a copy of eligible requests to a candidate model while the incumbent still determines the action. Compare outputs, latency and failure rate without sending the candidate’s decision downstream. Shadow traffic can still incur compute and data-handling obligations, so keep access and retention controls. It does not prove how users will react to candidate actions.

Make canary scope explicit

Route a small, stable slice of live traffic to the candidate. Use a deterministic assignment key so one customer is not switched between models on each request. Watch latency, error rate, abstention and downstream decision outcomes. If labels arrive late, do not pretend the first hour’s proxy metrics establish final quality. Label maturity changes the evaluation clock.

Define stop and rollback before launch

Set quantitative guardrails and an owner for the rollout. Keep the old artifact, preprocessing and class mapping loaded or rapidly recoverable. On a guardrail breach, stop new candidate assignments and restore the prior pointer. Record in-flight requests and whether decisions already sent downstream require repair.

Test routing and failure

Use a fixed key to confirm assignment stability, then inject a candidate timeout and incorrect class mapping. The timeout must trigger the documented fallback; the mapping failure should stop rollout before it reaches the canary. Compare candidate and incumbent over the same request population, noting where eligibility rules differ.

Implementation

python
import hashlib

def assign_canary(customer_id, rollout_percent):
    if not 0 <= rollout_percent <= 100:
        raise ValueError("invalid rollout percentage")
    fingerprint = hashlib.sha256(customer_id.encode("utf-8")).digest()
    bucket = int.from_bytes(fingerprint[:4], "big") % 10_000
    return bucket < rollout_percent * 100

Performance and operating cost

Shadowing can nearly double model inference work for mirrored requests; a canary adds little per-request compute but requires parallel artifact storage and monitoring. Hash assignment is O(length of key).

Common Mistakes

  • Do not let shadow outputs perform real actions.
  • Do not randomize assignment independently for each request.
  • Do not discard the previous preprocessing artifact before rollback is safe.

Read next

Continue the workflow: Forecast monitoring: separate data delay, demand shift and model failure.

Continue the workflow: Inference latency budgets: measure queue, feature and model time.

Continue the workflow: Project: promote a receipt model with artifact and rollback evidence.

Continue the workflow: Model incidents: build a release and evidence timeline before rollback.

Continue the workflow: Model experiments: separate assignment from actual exposure.

Continue the workflow: Hosted-model migration: paired replay and reversible cutover.

ai-data
mlops
Storage details