A dataset-intake prompt should establish what one row represents, which time window and population were exported, the meaning and unit of each relevant column, and how null, duplicate, or corrected records are represented. It should ask for unknown definitions instead of inferring them from a few sample rows. Keep the schema and source revision alongside the analysis request so a later report can be reproduced. A model can flag suspicious columns, but the data owner must confirm business semantics. No summary is reliable if a field named total alternates between cents and whole currency units.
Dataset intake prompts: define a column before analyzing it
Operational case
An operations analyst receives a shipment export with 4,823 rows. The column arrival_time contains UTC timestamps for most rows but local depot time for a small legacy slice. A status value of closed may mean delivered or canceled depending on the carrier version. The assistant cannot compute a late-delivery rate from this sample until the time zone and status mapping are supplied. It produces a compact intake question list and a record-level exception bucket. The analyst then confirms that 100 rows are exact duplicates, leaving 4,723 unique shipments; 23 still lack a resolved status, so the known-status cohort has 4,700 shipments.
Dataset: shipment_export_v7; row grain: one shipment event.
Required: shipment_id, promised_at, delivered_at, status, carrier_version.
Unknown: legacy timestamp zone; meaning of closed by carrier.
Raw=4,823; exact duplicates=100; unique=4,723; unknown=23.
Known-status denominator=4,700 after time and status review.Performance and operating cost
Scanning R rows for missing fields and duplicate keys takes O(R) expected time with a hash map and O(U) space for U unique keys. Time-zone conversion and versioned status mapping add work proportional to affected rows. A model call is cheap compared with publishing a wrong denominator, but repeated whole-file prompts waste tokens and may expose private records. Supply a typed schema, aggregate profile, and selected exceptions first. Measure unresolved definitions as a first-class result; do not bury them in a footnote after a confident chart.
Common Mistakes
- Do not infer column units from a small sample.
- Do not treat an event row count as a shipment count without checking duplicates.
- Do not mix local and UTC timestamps in one comparison.
Connected lessons
- Prompt engineering applications
- Prompt Engineering
- Prompt inputs: normalize records before asking for conclusions
- Output contracts: parse a result and preserve an explicit unknown state
- Retrieved evidence: reconcile versions and conflicting facts
- Classification prompts: write the label boundary first
- Aggregation prompts: pin the denominator and recompute the rate
- Report summaries: preserve contrary results and missing data
- Chart prompts: inspect axes, units, and missing series first
- Project: verify a shipment performance report
- Prompt engineering for data workflows
Continue with: Workbook prompts: name sheets, ranges, types, and provenance.
Continue with: Chart prompts: state the decision and measured quantity.
Continue with: Interview prompts: write a codebook with inclusion rules.
Continue with: Forecast prompts: define target, horizon, and time grain.
Continue with: Survey prompts: validate answer options and skip paths.
Continue with: Analytics prompts: define the metric and event contract.
