Segmenting an experiment after a surprising result can help diagnose a mechanism, but every extra slice offers another chance to see noise. A prompt should distinguish prespecified subgroups from exploratory investigations, keep denominators and assignment rules on both sides, and record whether the slice was chosen before or after outcome review. A guardrail movement deserves its own validity and practical-impact check. Do not dismiss a possible harm solely because the primary metric improved; the joint release decision must consider both.
Experiment prompts: investigate slices and guardrail movement
Operational case
Aster has 96 support tickets among 4,800 control users, or 2%, and 144 among 4,800 treatment users, or 3%. The observed guardrail difference is one percentage point. The analyst checks ticket taxonomy, duplicate reports, attribution window, and whether form confusion could explain the movement. A model may suggest a mobile slice, but unless it was planned, that slice is exploratory and cannot rescue a blocked rollout by itself.
Support guardrail: control 96/4,800 = 2%
Treatment 144/4,800 = 3%
Observed difference: +1 percentage point
Investigate taxonomy, duplicates, window, and mechanism
Post hoc slices: exploratory labels retainedPerformance and review cost
Scanning N ticket records and assignment IDs is O(N) expected time with indexed joins; K subgroup analyses can multiply compute and review costs. The more important cost is interpretive: many unplanned slices make one attractive story easy to find. Keep a slice ledger and require a separate confirmation plan for a new subgroup hypothesis.
Common Mistakes
- Do not ignore the higher ticket rate because completion increased.
- Do not promote a post hoc subgroup story to the primary result.
- Do not compare ticket counts without the assigned-user denominators.
Connected lessons
- Production prompt engineering
- Prompt Engineering
- Analytics prompts: explain segment gaps without inventing causes
- Aggregation prompts: pin the denominator and recompute the rate
- Experiment prompts: fix hypothesis, unit, and exposure
- Experiment prompts: investigate assignment and exposure imbalance
- Experiment prompts: freeze the primary metric and guardrails
- Experiment prompts: reconcile missing outcomes by assigned unit
- Experiment prompts: separate effect size from uncertainty
- Experiment prompts: gate rollout on a reproducible packet
- Project: review Aster's signup experiment
- Product experiment prompt decisions
