A prompt evaluation gate is a repeatable check applied to the exact prompt, model settings, tool schema, and evidence contract proposed for release. Put fast deterministic checks first: parse templates, verify required variables, validate output schemas, and execute local validators. Then replay a frozen labeled set for changes that can alter model behavior. CI should record the candidate revision and case-set version; a green run for an earlier branch is not proof about the current merge candidate. Stochastic outputs require raw case results and stable thresholds, not a claim that every run will produce identical text.
Prompt changes in CI: test the merge candidate
Decision in practice
A claims team edits a prompt to make approval explanations shorter. The template renders, but a high-impact held-out case changes from review to approve when the receipt is absent. The CI job blocks that merge despite an improved average length score. A second change only corrects a typo in a non-executed comment; the team can run static checks and a small smoke set before a scheduled full replay, according to its documented policy. A reviewer inspects the failing case, prompt diff, model setting, and evidence snapshot. The job does not silently replace the failed case with a more favorable one.
Merge candidate: prompt revision PE-118, schema V4, model setting MS-7.
Static gates: variables present; schema parses; tool names match catalog.
Replay: frozen set EV-47, including missing receipt and wrong-account cases.
Block: any critical false approval or unauthorized effect.
Report: per-case baseline and candidate output, slice counts, latency, cost.
On failure: preserve artifact and case IDs for review.Performance and operating cost
Static validation is typically linear in template and schema size. A replay over N cases and V prompt variants requires roughly O(NV) model calls; large sets may run in a separate job when the release policy permits, but high-impact regression cases belong on the merge gate. Cache only when input, model, and settings match exactly. Flaky model judgments should be adjudicated rather than automatically rerun until green. Count dollars, elapsed time, and human review together with error rates; a cheap gate that misses a critical false approval is not economical.
Common Mistakes
- Do not test a different commit from the one being merged.
- Do not allow aggregate improvement to hide a critical case failure.
- Do not rerun a flaky grader until it produces a passing result.
Connected lessons
- Production prompt engineering
- Prompt Engineering
- Code lab: block a critical prompt regression
- Evaluation leakage: keep the release test independent
- Continuous integration: test the merge candidate
- Git change control: small merges and protected branches
- Prompt release artifacts: version the whole decision path
- Prompt telemetry: measure failures without copying private payloads
- Project: measure a retrieval-backed answer gate
Continue with: Paired prompt evaluation: count changes, then inspect uncertainty.
Field guide: Prompt release review: assemble the decision packet.
Continue with: Project: review a checkout release from diff to incident.
Continue with: Retrieval release: test permission changes and deletion.
Continue with: Project: build a North Pier recurring review.
Continue with: Project: verify a Manifest Gate importer fix.
