Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Log templates: constants, variables and event identity

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A log line mixes stable event wording with values that change per occurrence. Parse both without discarding the event that the line reports.

Define an event before masking values

“Gateway retry rejected for request 47” contains a stable action and a request identifier. A template may represent the line as “Gateway retry rejected for request {request_id}.” Do not replace every number blindly: HTTP 503 or a protocol version can be part of the event meaning. Record source service, timestamp, severity, raw line, template ID and parser version. Quantity parsing handles measured values; event-template parsing decides which tokens vary across instances.

Preserve parameter roles

Two templates with the same words but different parameter roles are not safely interchangeable. “Worker 47 failed on node 82” and “Worker 82 failed on node 47” have different actors. Give each parameter a name, type, source span and validation rule. Avoid a single undifferentiated wildcard list that loses actor and location. Predicate roles help separate who failed from where it happened, especially when templates are learned from noisy logs.

Choose a safe unknown state

A new release may change wording, add a field or introduce a new error code. If no known template matches, keep the raw line and report an unparsed state. Do not force it into the nearest template merely to keep the dashboard tidy. A parser can also match several templates; retain those candidates for review. Version the template registry alongside deployment and source-code revisions, since a previously correct pattern may become misleading after a logger changes.

Evaluate downstream incidents

Measure exact template assignment, parameter extraction, unknown rate and event-order errors by service and release. Include log lines that differ only by negation, error code or parameter order. A high compression ratio can hide harmful merging of distinct events. The log triage project tests whether parsed events support an incident timeline, not just whether the parser groups similar strings.

Implementation

python
import re

RETRY_REJECTED = re.compile(
    r"Gateway retry rejected for request (?P<request_id>[0-9]+)")

def parse_gateway_line(raw_line, service, revision):
    match = RETRY_REJECTED.fullmatch(raw_line)
    if match is None:
        return {"state": "unparsed", "raw_line": raw_line,
                "service": service, "source_revision": revision}
    return {"state": "parsed", "template_id": "gateway-retry-rejected-r3",
            "request_id": match["request_id"], "service": service,
            "source_revision": revision, "raw_line": raw_line}

assert parse_gateway_line("Gateway retry rejected for request 47",
                          "gateway-west", "deploy-r5")["request_id"] == "47"
assert parse_gateway_line("Gateway retry accepted for request 47",
                          "gateway-west", "deploy-r5")["state"] == "unparsed"

Performance and operating cost

Matching one anchored, fixed-shape line is O(n) in its length n for this pattern. A registry with t candidate patterns can cost O(t × n) without an index, and poorly designed expressions may behave much worse. Keep unknown lines for review; silently misclassifying a critical event can cost more than storing unparsed text.

Common Mistakes

  • Masking every number, including a meaningful error code.
  • Dropping the raw line after a template match.
  • Forcing unknown wording into the closest known template.
  • Evaluating only line compression instead of event correctness.

Read next

ai-data
natural-language-processing
Storage details