A study-extraction prompt should produce a record, not a free-form conclusion. Capture study ID, report version, design, allocation method if stated, eligible population, outcome definition, time point, numerator, denominator, missing outcomes, and location of the source result. Mark unavailable fields as unknown. If two documents describe one study, link them to one study ID instead of counting them as independent evidence. Keep the original wording or a bounded excerpt in the private review packet where permitted, then write a fresh synthesis from verified fields. Page or table anchors let a reviewer check each number before any calculation or claim leaves the workflow.
Study extraction prompts: preserve the result behind each claim
Operational case
ST-47's abstract reports a 7.5 percentage-point reduction, while its results table supplies 48 misroutes among 240 control parcels and 30 among 240 candidate parcels. The extraction packet records both counts and the results-table anchor. A later conference summary repeats ST-47's same cohort and does not become a third trial. ST-88 reports 18 of 120 in control and 21 of 120 with the candidate panel. Its adverse result remains in the record. ST-61 lacks a stable allocation mechanism and has a different route mix; its design field stays observational rather than being promoted to a randomized trial because the model expects one.
ST-47 | controlled line trial | control 48/240 | candidate 30/240 | full shift.
ST-88 | controlled line trial | control 18/120 | candidate 21/120 | full shift.
ST-61 | changed-shift observation | route mix differs | context only.
Duplicate report: ST-47 conference summary -> same cohort, one study ID.
Required: result anchor, missing-outcome field, report version.Performance and operating cost
For N reports, extraction is O(N) model jobs plus manual review of each uncertain field. Deduplicating reports by study identity needs a maintained index; title similarity alone can merge different cohorts or split one cohort into several entries. Keeping explicit numerators and denominators costs more text than a headline percentage, but makes arithmetic and missingness checkable. If a report only states a percentage without a denominator, do not manufacture counts. Anchor verification is a human or parser step, not something a self-assured model statement can replace.
Common Mistakes
- Do not count two reports of one cohort as two studies.
- Do not fill an unstated denominator from a nearby chart or abstract.
- Do not let an abstract headline override the checkable result table without review.
Connected lessons
- Prompt engineering applications
- Prompt Engineering
- Document prompts: anchor each field to a page and resolve conflicts
- Dataset intake prompts: define a column before analyzing it
- Evidence IDs: make generated claims auditable against supplied records
- Research prompts: fix the question and eligibility first
- Study effects: calculate the measure outside the model
- Conflicting studies: inspect design before averaging outcomes
- Research synthesis release: expose the method and its limits
- Project: synthesize conflicting parcel-scan trials
- Research-synthesis prompt decisions
