A change-failure numerator counts production deployments that cause a problem requiring immediate intervention, under a documented attribution window and service boundary. A build that fails before release is a pipeline failure, not a production change failure. A rollback, urgent hotfix, or manual mitigation can establish intervention, but the team still needs evidence that a specific deployment caused the impairment. One incident can involve several deployments; count each implicated deployment once rather than counting alert notifications. Attribution can be uncertain when a dependency fails at the same time. Keep an unknown state until review rather than forcing every incident into a neat numerator.
Change fail rate: attribute immediate intervention to a production deployment
Operational decision
During a month, the portal has 26 successful production activations. Three distinct activations require urgent intervention after user-path checks fail. The observed change fail rate is 3 divided by 26, or about 11.5 percent. Two alerts for the same failed activation do not make it two failures. A database outage that begins before a deploy is investigated separately and is not automatically assigned to the newest version. The release ledger stores incident ID, first user impact, intervention type, evidence, attribution confidence, and review owner. Recalculate a past period when a review changes attribution, but keep a visible revision of the report.
Reporting period: 26 active production deployments
Attributable failures: 3 distinct deployment IDs
Change fail rate: 3 / 26 = 11.5% after rounding
Duplicate alerts: deduplicated by deployment and incident
Uncertain cause: reviewed separately, not guessed
Interventions: rollback, hotfix, or urgent mitigationCost and verification
A join between D deployments and I incident links is manageable with indexed IDs, but human attribution is the costly part. Keep the counting rule short enough for an operator to apply during an incident. A low rate can be misleading if teams stop recording hotfixes or use a long detection delay that pushes failures outside the chosen window. Read the rate with deployment volume, user impact, and SLO results; a team can achieve zero failures by shipping nothing.
Common Mistakes
- Do not count pre-production build failures in the production change-failure numerator.
- Do not multiply one incident by its alert count.
- Do not equate the newest deployment with the cause before checking evidence.
Connected lessons
- DevOps: delivery, infrastructure, and reliable operations
- Deployment frequency: count activated production changes from an event ledger
- Incident response: contain impact, then learn
- Canary analysis: compare a small cohort without hiding its failures
