An online prompt experiment compares a candidate with a baseline on live traffic using a stable assignment unit, eligibility rule, and predeclared outcome measures. Offline checks come first. If the workflow can cause an external effect, the experiment must keep the same application permissions and independent validators for both variants. Assign a user or account consistently so one party does not alternate between policies mid-task. Measure completion, correction, escalation, latency, and cost together with critical error counts. Define a stop rule and an operator before exposure; do not inspect the dashboard repeatedly and declare success at the first favorable fluctuation.
Online prompt experiments: define exposure and stop rules first
Decision in practice
A support assistant proposes a new clarification prompt for account changes. The team first passes offline wrong-account and missing-authorization cases. It then sends a small eligible slice of read-only support traffic to the candidate, assigned by account ID. The experiment excludes payment and cancellation actions. The candidate reduces unnecessary questions, but two cases omit the final verification step. The predeclared stop rule pauses exposure and sends those cases to review. The team checks assignment, trace IDs, and event definitions before blaming the wording alone.
Assignment unit: account_id; eligibility: read-only support.
Variants: baseline PEB-118, candidate PEB-124.
Primary: correct resolution after review; guardrail: omitted verification.
Stop: any confirmed wrong-account action or repeated missing step.
Window: fixed review period; owner: support on-call.
Record: assignment, bundle ID, validator result, correction outcome.Performance and operating cost
Routing and recording a fixed set of fields add O(1) work per eligible request; analysis over M events is O(M). The main cost is enough traffic and human review to distinguish a real improvement from chance and case-mix changes. A small canary may detect severe failures but rarely estimates a subtle quality gain precisely. Early stopping for harm is appropriate; early stopping for a favorable score can inflate the apparent improvement. Keep assignment and exclusion rules fixed, and do not use downstream actions as a proxy when their logs are incomplete.
Common Mistakes
- Do not expose a candidate that failed its offline critical gate.
- Do not randomize each turn when account-level consistency matters.
- Do not choose the experiment's winning metric after seeing the dashboard.
Connected lessons
- Production prompt engineering
- Prompt Engineering
- Prompt releases: version the whole decision path and keep a rollback
- Prompt changes in CI: test the merge candidate
- Progressive delivery: canary checks and rollback
- Paired prompt evaluation: count changes, then inspect uncertainty
- Evaluation labels: adjudicate disagreement before scoring a release
- Adversarial case mutations: test the boundary, not a magic phrase
- Evaluation case ledgers: revise labels without erasing history
- Project: govern a claims-assistant evaluation board
- Prompt evaluation governance decisions
Field guide: Prompt release review: assemble the decision packet.
Continue with: Experiment prompts: fix hypothesis, unit, and exposure.
Continue with: Search prompts: interpret clicks and gate a ranking release.
