A canary release sends a limited share of traffic to a new revision while the older revision still serves most users. Its safety depends on comparing meaningful outcomes across those revisions. A tiny cohort may not include the operation that changed; a large cohort may expose too many users before a fault is detected.
Canary analysis: compare a small cohort without hiding its failures
Operational decision
A parcel pricing release starts with internal users and then five percent of eligible requests. Pin the same request types and regions in both old and new samples, or record where they differ. The policy record defines what the release controller and operator should measure; it is not a controller API. Compare error rate, p95 latency, and a business check such as duplicate prices. Wait long enough for delayed queue effects to appear before promotion. If the stop threshold is crossed, return new traffic to the old revision while checking whether side effects or schema changes need separate recovery. A feature flag can stop a code path, but it is not a substitute for measuring the actual served revision. Record cohort size and confidence limits so a clean result from seventeen requests is not called strong evidence.
Parcel pricing canary
Exposure: internal cohort, then 5 percent of eligible traffic
Compare: same route, region, and request mix
Stop: error ratio above 1.8 percent or p95 above 740 ms
Check: duplicate price events and delayed queue effects
Promote: named decision after observation window
Recover: route old revision and inspect completed writesCost and verification
Running two revisions consumes extra compute and monitoring effort. A longer observation window improves evidence but slows delivery; a shorter one can miss rare or delayed defects. Metrics tagged by revision increase cardinality, so limit tags to stable identifiers. Evaluate both absolute user impact and relative comparison: a canary can look better than an already failing baseline while still violating the service objective. Preserve the decision record with the release digest.
Common Mistakes
- Do not promote on a small, unrepresentative sample.
- Do not compare different routes or regions without noting the bias.
- Do not call traffic rollback a reversal of completed writes.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- Progressive delivery: canary checks and rollback
- Feature flags: stop exposure without pretending code vanished
- Release evidence: tie one deployed digest to one approval decision
- SLOs and error budgets: turn reliability into a decision
