Build a disposable event ledger for one service. Seed 43 pipeline runs, 29 production attempts, and 26 distinct activations across a month. Attach artifact digests and commit sets to each activation, then create three deployment-caused incidents and four unplanned repair activations. The exercise passes only when repeated events, previews, failed pre-activation attempts, and unrelated outages are classified without inflating the headline measures. Use synthetic event IDs and avoid importing employee-performance data.
Project: build an auditable delivery and recovery measurement ledger
Define the event contract
Store immutable deployment ID, service, environment, source revision, artifact digest, attempt result, activation time, and release intent. Keep commit-to-artifact membership in a separate relation. Link incidents to deployment IDs with impact start, detection, intervention, verified restoration, and attribution confidence. Require a reason for unknown or corrected classifications. Run a monthly report with a half-open UTC interval so midnight events cannot fall in two months. Review the first and last day manually against source receipts.
Synthetic acceptance
Production activations: 26, not 43 pipeline runs
Pre-activation failures: 3, retained outside frequency numerator
Attributable failed activations: 3 / 26 = 11.5%
Unplanned repair activations: 4 / 26 = 15.4%
First impact 14:06 to verified recovery 14:55 = 49 min
Unknown attribution: visible queue for reviewInject ledger faults
Replay one deployment event with the same ID and confirm the count remains 26. Promote an earlier artifact after rollback and prove that its original commits retain their first-live timestamps. Add two alerts for one failed activation and keep the change-failure numerator at one for that activation. Move a failed attempt to the end of the month; it must not become a successful activation merely because the pipeline ended. Make one incident cause two repair deployments and confirm both count as rework while the incident remains one record.
Review the measures together
Show deployment timeline, change lead-time distribution, change fail rate, failed-deployment recovery durations, rework share, and missing-link counts. Compare the same service across periods without changing definitions. A shorter lead time is not a victory if user impact grows; a zero failure rate is not a victory if no changes reach production. Preserve report revisions so a later incident review can correct attribution without erasing the first published result.
Common Mistakes
- Do not equate pipeline runs with live deployments.
- Do not infer causation solely from event timing.
- Do not stop recovery time at the rollback command.
Connected lessons
- Deployment frequency: count activated production changes from an event ledger
- Change lead time: join committed work to its first live artifact
- Change fail rate: attribute immediate intervention to a production deployment
- Failed deployment recovery time: keep impact, detection, and restoration clocks distinct
- Deployment rework: identify unplanned repair releases without hiding planned work
- DevOps projects
