Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Project: release an incident-runbook search service

Last updated: 7 Oct 20265 min read
project
AdvancedBy AITrove Editorial

Build a search endpoint that finds a permitted runbook revision, explains ranking failures and refuses to claim an answer when the collection has none.

The result contract

An on-call engineer searches for a failure symptom and receives a ranked list of passages with document title, revision, owner, last-reviewed time and permission-safe link. Return a no-result state if nothing passes relevance and authorization. A search hit is not an instruction to execute a fix. Preserve the exact query and result IDs for audit, but avoid placing incident secrets into broad telemetry. Decide what freshness means before indexing: a retired runbook should never outrank an active replacement because it has more matching words.

Assemble the collection

Chunk at document structure boundaries so a passage retains its prerequisite and warning. Give each chunk a stable ID tied to document revision and offsets. Build lexical and dense indexes from the same authorized snapshot, using a versioned embedding contract. Create reviewed queries from incident tickets after the relevant runbook revision existed. Include literal error codes, paraphrases, stale revisions, no-answer queries and restricted documents. Group duplicated incidents in one split.

Rank, filter and test

Retrieve bounded lexical and dense candidates, enforce access rights, fuse their ranks and rerank the small merged list. Run permission tests with several user roles, including a user who recently lost access. Compute recall before and after filtering; do not hide a poor index behind a strong reranker. Set release gates for useful-hit rank, no-answer precision, authorization failures and latency. Hybrid ranking defines the decision sequence.

Roll over safely

Build the new index under a fresh version, run the same audit queries, and atomically move the alias when checks pass. Keep the previous index for rollback. Record every source deletion and permission change; a delayed index refresh must not expose revoked content. Watch for high no-result rates on a new product line and collect reviewed queries for the next update. The broader retrieval learning path covers generation after retrieval, but this service should work as search first.

Implementation

python
def release_gate(metrics):
    required = {"authorized": True, "candidate_recall_at_20": 0.88,
                "no_answer_precision": 0.81, "p95_ms": 430}
    missing = set(required) - set(metrics)
    if missing:
        raise ValueError(f"missing release metrics: {sorted(missing)}")
    return (metrics["authorized"] is True
            and metrics["candidate_recall_at_20"] >= required["candidate_recall_at_20"]
            and metrics["no_answer_precision"] >= required["no_answer_precision"]
            and metrics["p95_ms"] <= required["p95_ms"])

audit = {"authorized": True, "candidate_recall_at_20": 0.91,
         "no_answer_precision": 0.84, "p95_ms": 396}
assert release_gate(audit)

Performance and operating cost

Index storage scales with text postings plus O(n·d) vector values for n passages of dimension d. Query latency combines two retrieval calls, batched permission checks and bounded reranking; measure the slowest stage separately. A fixed candidate depth limits cost but may miss relevant text, so publish recall at that exact depth. Reindexing should happen off the serving alias to keep query latency stable.

Common Mistakes

  • Indexing old and new revisions without a freshness rule.
  • Treating a relevant but unauthorized runbook as a successful hit.
  • Moving an alias before no-answer and access tests pass.
  • Claiming a generated fix is grounded solely because search returned a passage.

Read next

Continue the workflow: Project: publish evidence-backed incident handoff summaries.

Continue the workflow: Project: answer runbook questions with source spans and abstention.

Continue the workflow: Project: build a reviewed runbook keyphrase index.

Continue the workflow: Project: release controlled runbook query reformulation.

Continue the workflow: Project: answer a local-language question from a foreign-language runbook.

ai-data
natural-language-processing
Storage details