Version a prompt, model and manual index together, reject unsupported instructions and stage a reversible summarizer release.
Project: release a maintenance-ticket summarizer with evidence gates
Freeze the bundle
A service team wants short summaries of equipment maintenance tickets. The application retrieves approved manual excerpts and reads ticket metadata through a limited tool. Pin the prompt revision, resolved model revision, manual-index snapshot, tool schema, output schema and fallback policy. Separate instructions from untrusted ticket text and retrieved excerpts. The manifest is the release unit; a prompt edit alone is not the complete change.
Build the evaluation packet
Assemble 500 cases with old and current manuals, conflicting notes, missing serial numbers, empty retrieval and a malicious sentence embedded in a ticket. Validate JSON fields and allowed enums mechanically. Have qualified reviewers check safety-relevant claims and whether the summary points to the evidence actually provided. When no manual passage supports an instruction, the system must abstain and send the ticket to a technician. The gate separates format, groundedness and fallback behavior.
Find the hidden regression
The candidate prompt produces shorter summaries and fewer formatting failures. Yet it uses a stale manual index that still contains an older maintenance interval. A global preference score rises while one unsupported instruction appears in a high-risk slice. Hold the bundle, rebuild the index from approved manuals and repeat the frozen test without editing the held-out cases. Check that the tool call remains within the ticket’s authorized scope. Privacy controls prohibit copying raw ticket text into unrestricted diagnostics.
Stage the corrected release
Deploy a small technician cohort, log bundle identity and validation status, and sample output review under the permitted retention policy. Watch abstention, manual escalations, unsupported claims, p99 latency and cost. A rollback restores the prior prompt, model, index and tool contract together. Publish the evaluation limits: five hundred cases and a clean reviewed sample are evidence for a controlled release, not a promise of error-free generation. Reversible promotion should name the exact prior bundle.
Implementation
def maintenance_summary_disposition(bundle, validated, review):
required = {"prompt", "model", "manual_index", "tool_schema"}
if any(not bundle.get(name) for name in required):
return "hold:bundle-incomplete"
if not validated["schema_ok"] or not validated["tool_scope_ok"]:
return "hold:contract"
if review["unsupported_instructions"]:
return "hold:unsupported-claim"
if not review["abstains_without_evidence"]:
return "hold:missing-abstention"
return "stage:technician-cohort"
bundle = {"prompt": "p47", "model": "text-r8",
"manual_index": "manuals-i31", "tool_schema": "ticket-t4"}
validated = {"schema_ok": True, "tool_scope_ok": True}
review = {"unsupported_instructions": 0, "abstains_without_evidence": True}
assert maintenance_summary_disposition(bundle, validated, review) == "stage:technician-cohort"
assert maintenance_summary_disposition(bundle, validated,
{**review, "unsupported_instructions": 1}) == "hold:unsupported-claim"
Performance and operating cost
The release decision is O(1) time and space, but evaluation requires generation, retrieval, tool calls and expert adjudication over many cases. Staged sampling adds review cost and reduces rollout speed. Those costs are preferable to a prompt change that silently pairs with stale manuals and produces an unsupported maintenance instruction.
Common Mistakes
- Shipping the new prompt against an old manual index.
- Using format success as a substitute for factual review.
- Letting ticket text request unauthorized tool actions.
- Rolling back only the prompt after discovering an index regression.
Read next
- Generative releases: bind prompt, model, tools and output contract
- Generative evaluation gates: grounded claims, schemas and abstention
- Inference logs: keep diagnostic joins without copying sensitive payloads
- Promotion evidence: bind evaluation, contract and rollback to one digest
- Test partial failure and deadline exhaustion in model chains
