A prompt release gate compares a candidate version against frozen cases and blocks it when a critical invariant fails, even if aggregate output looks better. Keep a baseline result for diagnosis and a labeled expected outcome for the gate. Run the same input and evidence versions for both candidates; otherwise the comparison mixes prompt effects with data changes. The gate below checks final labels on critical cases and reports any changed result for review.
Code lab: block a critical prompt regression
Decision in practice
A claims prompt revision fixes several routine explanations but changes CL-411 from deny to approve despite an expired receipt. The baseline and candidate are replayed against the same policy snapshot. The release is blocked because the critical case is wrong. A second changed case, CL-412, remains review and needs no final-decision override. The team records the candidate version, failing case, and evidence snapshot so the author can make a targeted repair before the next run.
frozen_cases = [
{"id": "CL-411", "critical": True, "expected": "deny", "baseline": "deny", "candidate": "approve"},
{"id": "CL-412", "critical": True, "expected": "review", "baseline": "review", "candidate": "review"},
{"id": "CL-413", "critical": False, "expected": "approve", "baseline": "review", "candidate": "approve"},
]
critical_failures = [
case["id"] for case in frozen_cases
if case["critical"] and case["candidate"] != case["expected"]
]
changed_cases = [
case["id"] for case in frozen_cases
if case["candidate"] != case["baseline"]
]
print("release:", "block" if critical_failures else "allow")
print("critical_failures:", ", ".join(critical_failures) or "none")
print("changed_cases:", ", ".join(changed_cases) or "none")
Expected output: release: block; critical_failures: CL-411; changed_cases: CL-411, CL-413
Performance and operating cost
A single replay comparison is O(N) time and O(N) worst-case space for the two result lists. Model calls, not list scans, dominate release cost; comparing both versions doubles calls when baseline outputs were not stored. Preserve sampled responses and model settings with the case snapshot. The critical gate is necessary but insufficient: check privacy, unauthorized tool effects, latency, and slice coverage separately. A release set that was used to write the prompt no longer gives an independent estimate.
Common Mistakes
- Do not promote on aggregate gains while a critical case fails.
- Do not compare different evidence snapshots as if only the prompt changed.
- Do not reuse a tuned development case as the sole release holdout.
Connected lessons
- Production prompt engineering
- Prompt Engineering
- Prompt releases: version the whole decision path and keep a rollback
- Evaluation leakage: keep the release test independent
- Code lab: reject unknown evidence IDs
- Code lab: score answered cases and abstentions
- Code lab: find pairwise judge order changes
- Prompt evaluation code labs
Related implementation
Continue with: Architecture release: require evidence, rollback, and ownership.
Continue with: Infrastructure release: verify live state and rollback limits.
Continue with: Localization prompts: gate release on catalog and rendered retests.
