A production incident is rarely a single exception. A case save may cross a browser, API, queue, and worker; each can report success while the final user action still fails. This track starts with the user-visible operation and follows its trace through asynchronous work. It sets measurable objectives, limits telemetry exposure and cardinality, and turns incident evidence into a containment or rollback decision. The goal is a reproducible diagnosis, not a dashboard filled with unrelated counters.
Topics in this track
- Trace Context Across Requests and Jobs — Follow one user action through HTTP and queued work without treating trace IDs as authorization.
- User Journey SLOs and Burn Alerts — Measure whether people can finish an operation and alert on meaningful error-budget loss.
- Telemetry Shapes, Redaction, and Cardinality — Keep event records useful for diagnosis while controlling data exposure and metric-series growth.
- Incident Containment and Evidence Timeline — Choose a reversible containment action from a coherent timeline of user impact and releases.
Prerequisite paths
Observability: connect user failure to a safe request trace; Release checks: prove the critical route and prepare a rollback; Telemetry Minimization and Retention.
Neighbor track
Cross-Tab and Offline Data Coordination.
Practice path
Build Project: case-save incident evidence and check decisions in Web Development: operations and sync decisions quiz.
Further connections
Abuse Decision Telemetry and False-Positive Rollback.
Further connections
Browser Performance Diagnosis and Measurement; Field Performance Observation and Sample Contracts.
