Turn raw gateway and worker logs into reviewable incident events without hiding unknown templates, privacy risks or release changes.
Project: parse incident logs with versioned event templates
Set the incident task
An operator needs to know whether retry rejection began before a rollback. Ingest timestamped log lines with service, environment and deployment revision. Parse a small approved template registry, store named parameters and keep unmatched raw lines in a review queue. Do not infer that a similar line means the same event. Template boundaries preserve constant event wording and parameter spans.
Build hard cases
Include logger revisions that change punctuation, reorder parameters, add an error code or rename a field while keeping the event text familiar. Add unknown events, copied log excerpts and request IDs that resemble sensitive tokens. Group all lines from an incident and deployment before splitting the evaluation set. Reviewers label exact template IDs, parameter roles and whether the event belongs in the incident timeline. A random line split would leak near-identical templates.
Stage parsing and chronology
Apply an anchored parser and parameter schema tied to the deployment revision. Route no-match and multi-match lines for review. Keep raw text restricted; publish only permitted event fields. Normalize timestamps under a declared zone policy and attach source line IDs to temporal edges. Parameter identity catches schema changes, while event-order checks detect impossible chronology.
Gate the release
Report wrong event merges, missed unknowns, parameter errors, privacy leaks and timeline reversals. Compare event counts by template and deployment so a parser revision cannot masquerade as recovery. Shadow the pipeline against resolved incidents, then keep the old template registry for rollback. A compact dashboard is useful only if an operator can open the restricted raw evidence for each displayed event under the right access role.
Implementation
def stage_log_event(parsed, permitted_fields, active_revisions):
if parsed["source_revision"] not in active_revisions:
return {"state": "review", "reason": "unknown-release"}
if parsed["state"] != "parsed":
return {"state": "review", "reason": "unparsed-line"}
safe = {field: parsed[field] for field in permitted_fields if field in parsed}
safe["state"] = "proposed-event"
return safe
parsed = {"state": "parsed", "template_id": "retry-rejected-r3",
"source_revision": "deploy-r5", "request_token": "secret-47"}
event = stage_log_event(parsed, {"template_id", "source_revision"}, {"deploy-r5"})
assert event == {"state": "proposed-event", "template_id": "retry-rejected-r3",
"source_revision": "deploy-r5"}
assert "request_token" not in event
Performance and operating cost
Checking a staged event and copying f permitted fields is O(f) expected time and space. Parsing all n log lines costs at least O(total input characters), with extra review for unknown templates. Keep raw logs under their own access and retention policy; a safe derived event must not imply that the upstream raw store is safe.
Common Mistakes
- Treating unmatched lines as harmless noise.
- Publishing raw request tokens through a derived event.
- Comparing event counts across releases without template-version context.
- Assigning a timeline position from processing order instead of event time.
