A treatment-control difference is a descriptive estimate. An experiment claim also needs the planned uncertainty method, randomization and data-quality checks, observation cutoff, and handling of repeated looks or many tested outcomes. Have code calculate the effect and its uncertainty from a locked dataset; the model should explain the result without fabricating a confidence interval or p-value. If the team repeatedly checks results and stops at the first favorable number, an ordinary one-time test no longer has its advertised error behavior. Preserve the planned stopping rule and all reviewed looks.
Experiment prompts: separate effect size from uncertainty
Operational case
Aster's assigned-user rates are 2,304/4,800 = 48% for control and 2,496/4,800 = 52% for treatment. The observed difference is four percentage points. The packet does not include a checked interval or a completed exposure-log investigation, so the model may report the descriptive numbers but cannot claim a verified lift. Aster also records when interim reviews occurred; it does not present the most favorable day as the final planned readout.
Control 2,304 / 4,800 = 48%
Treatment 2,496 / 4,800 = 52%
Observed difference = +4 percentage points
Checked interval or test result: absent
Inference claim: withheldPerformance and review cost
Computing two rates from validated counters is O(1); preparing the per-user table is O(N) for N assigned users. Uncertainty calculations depend on the prespecified design and may be more complex when units are clustered or outcomes are correlated. The prompt should keep the arithmetic and inference artifacts separate so a reviewer can reproduce each one.
Common Mistakes
- Do not invent a confidence interval from an aggregate summary.
- Do not equate four percentage points with statistical certainty.
- Do not hide repeated interim looks.
Connected lessons
- Production prompt engineering
- Prompt Engineering
- Study effects: calculate the measure outside the model
- Online prompt experiments: define exposure and stop rules first
- Experiment prompts: fix hypothesis, unit, and exposure
- Experiment prompts: investigate assignment and exposure imbalance
- Experiment prompts: freeze the primary metric and guardrails
- Experiment prompts: reconcile missing outcomes by assigned unit
- Experiment prompts: investigate slices and guardrail movement
- Experiment prompts: gate rollout on a reproducible packet
- Project: review Aster's signup experiment
- Product experiment prompt decisions
