A release stages candidate observation, limited live traffic and rollback against explicit health and decision metrics.
Shadow and canary rollout: compare a candidate without losing a rollback
Use shadow for observation
Send a copy of eligible requests to a candidate model while the incumbent still determines the action. Compare outputs, latency and failure rate without sending the candidate’s decision downstream. Shadow traffic can still incur compute and data-handling obligations, so keep access and retention controls. It does not prove how users will react to candidate actions.
Make canary scope explicit
Route a small, stable slice of live traffic to the candidate. Use a deterministic assignment key so one customer is not switched between models on each request. Watch latency, error rate, abstention and downstream decision outcomes. If labels arrive late, do not pretend the first hour’s proxy metrics establish final quality. Label maturity changes the evaluation clock.
Define stop and rollback before launch
Set quantitative guardrails and an owner for the rollout. Keep the old artifact, preprocessing and class mapping loaded or rapidly recoverable. On a guardrail breach, stop new candidate assignments and restore the prior pointer. Record in-flight requests and whether decisions already sent downstream require repair.
Test routing and failure
Use a fixed key to confirm assignment stability, then inject a candidate timeout and incorrect class mapping. The timeout must trigger the documented fallback; the mapping failure should stop rollout before it reaches the canary. Compare candidate and incumbent over the same request population, noting where eligibility rules differ.
Implementation
import hashlib
def assign_canary(customer_id, rollout_percent):
if not 0 <= rollout_percent <= 100:
raise ValueError("invalid rollout percentage")
fingerprint = hashlib.sha256(customer_id.encode("utf-8")).digest()
bucket = int.from_bytes(fingerprint[:4], "big") % 10_000
return bucket < rollout_percent * 100Performance and operating cost
Shadowing can nearly double model inference work for mirrored requests; a canary adds little per-request compute but requires parallel artifact storage and monitoring. Hash assignment is O(length of key).
Common Mistakes
- Do not let shadow outputs perform real actions.
- Do not randomize assignment independently for each request.
- Do not discard the previous preprocessing artifact before rollback is safe.
Read next
- Model promotion: require evidence before changing the serving pointer
- Model monitoring: separate input drift, data faults and delayed outcomes
- Inference logs: keep diagnostic joins without copying sensitive payloads
- Project: release receipt triage with lineage, canary checks and rollback
Continue the workflow: Forecast monitoring: separate data delay, demand shift and model failure.
Continue the workflow: Inference latency budgets: measure queue, feature and model time.
Continue the workflow: Project: promote a receipt model with artifact and rollback evidence.
Continue the workflow: Model incidents: build a release and evidence timeline before rollback.
Continue the workflow: Model experiments: separate assignment from actual exposure.
Continue the workflow: Hosted-model migration: paired replay and reversible cutover.
