When a web request fails, the user needs a clear recovery path and the team needs enough context to find the server event. A correlation ID can connect a client-visible error to a protected server log entry. Logs should capture the route pattern, status, duration, and failure class; avoid raw session cookies, tokens, passwords, or unnecessary submitted notes. Aggregate rates and latency trends reveal problems that one smoke check misses. The example produces a structured event from known safe fields. A real service should generate or validate request IDs at its edge and pass them through downstream work, while keeping the full diagnostic trace behind access controls.
Observability: connect user failure to a safe request trace
Case study
A reviewer sees 'Save unavailable' after submitting a note. The page includes request ID 'r-47-29' in the support detail, while the server log has the same ID, route pattern, 503 status, and 218-millisecond duration. The log omits the note and cookie. An operator can see that the case write route failed across many requests after a release, trigger rollback, and then inspect the underlying database error. A 422 validation error should not page an operator like a 503 outage; classification and aggregation matter as much as the log line itself.
Working contract
function requestEvent({ requestId, route, status, durationMs }) {
return { requestId, route, status, durationMs,
failureClass: status >= 500 ? "server" : status >= 400 ? "client" : "none" };
}
const event = requestEvent({
requestId: "r-47-29", route: "/cases/:caseId/notes", status: 503, durationMs: 218
});
console.log(JSON.stringify(event));Observed output
{"requestId":"r-47-29","route":"/cases/:caseId/notes","status":503,"durationMs":218,"failureClass":"server"}Cost and tradeoffs
Writing one fixed-size event is O(1) application work and storage per request, but high traffic makes total log volume O(R) for R requests. Capturing full bodies would multiply storage and privacy risk; bounded fields are cheaper and safer. Histograms or rate counters add aggregation cost but make trends easier to spot. Sampling can reduce routine trace volume, yet security and error events may need a different retention rule. Measure alert quality: a noisy alert that no one trusts can hide the critical failure it was meant to expose.
Common Mistakes
- Do not log session cookies or raw anti-forgery tokens.
- Do not make a 422 user correction look like a server outage.
- Do not expose internal exception details in a public error page.
- Do not rely on one successful smoke check as ongoing monitoring.
Continue through the stack
Release checks: prove the critical route and prepare a rollback; Tests across boundaries: assert behavior, not implementation text; DevOps: delivery, infrastructure, and reliable operations.
Failure trace
A write returns 500 for one reviewer, but the logs record only 'request failed' with no request ID, route, or stage. The operator cannot tell whether validation, database commit, or notification delivery failed; retrying blindly may duplicate work. Emit a safe correlation ID and structured stage outcome, then expose an error state that tells the user whether a retry is safe.
Verification
- Trace one request ID from browser error to server route and storage outcome.
- Force a post-commit notification failure and confirm the record and pending work are visible separately.
- Check that logs omit note text, credentials, and other sensitive payloads.
Decision note
More logging is not the goal. Record enough to distinguish failure classes, bound retention, and connect alerts to a recovery action that an operator can actually perform.
Further connections
Telemetry Minimization and Retention.
Further connections
Production Signals and Incident Decisions; Trace Context Across Requests and Jobs; User Journey SLOs and Burn Alerts.
