A prompt for conflicting studies should ask why estimates differ before asking for one verdict. Check allocation, baseline differences, missing outcomes, outcome definitions, follow-up windows, and whether several reports describe the same participants. Separate measured observations from explanations for heterogeneity. An observational comparison can be useful context, yet a changing route mix may confound any before-after difference. A model should not assign a numeric bias score from thin report text or dismiss a contrary result as an outlier because it is inconvenient. A qualified reviewer chooses whether quantitative pooling is sensible; absent that plan, show the study-level outcomes and the limits of comparison.
Conflicting studies: inspect design before averaging outcomes
Operational case
ST-47 favors the new panel by 7.5 points; ST-88 goes the other direction by 2.5 points. The sites used different parcel mixes, but the packet does not say whether that explains the difference. ST-61 appears favorable in a changed-shift comparison, yet the later shift handled fewer oversized parcels. The assistant must not let ST-61 break the tie or declare two favorable 'studies' versus one unfavorable trial. It labels ST-61 observational, keeps the two controlled trials distinct, and requests allocation details, missing-outcome counts, and route composition before a stronger synthesis.
ST-47: candidate lower by 7.5 points; controlled line trial.
ST-88: candidate higher by 2.5 points; controlled line trial.
ST-61: later shift differed in route mix; observational context.
Unresolved: allocation details, missing outcomes, site comparability.
Do not pool or vote until comparability and analysis plan are reviewed.Performance and operating cost
Comparing K studies across D predeclared design fields costs O(KD) reviewer checks, even before statistical analysis. A single aggregate sentence is cheaper to write but can conceal missing outcomes or an incompatible outcome definition. Build a structured comparison record once, then reuse it for an editorial summary and a reviewer queue. If source documents disagree about a study's denominator, resolve that conflict at the report level before calculating. Model generation time is rarely the limiting cost; locating and adjudicating incomplete methods is.
Common Mistakes
- Do not count studies as votes without considering design and cohort overlap.
- Do not call an observational shift comparison a controlled trial.
- Do not invent a reason why two valid estimates disagree.
Connected lessons
- Prompt engineering applications
- Prompt Engineering
- Retrieved evidence: reconcile versions and conflicting facts
- Cross-modal evidence: keep conflicting observations separate
- Paired prompt evaluation: count changes, then inspect uncertainty
- Research prompts: fix the question and eligibility first
- Study extraction prompts: preserve the result behind each claim
- Study effects: calculate the measure outside the model
- Research synthesis release: expose the method and its limits
- Project: synthesize conflicting parcel-scan trials
- Research-synthesis prompt decisions
