Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Exploratory analysis without peeking: inspect the data and preserve the test

Last updated: 5 Oct 20265 min read
tutorial
BeginnerBy AITrove Editorial

Exploratory data analysis finds distributions, impossible values and unstable relationships while keeping the final evaluation set out of decisions.

Begin with contracts, not plots

Write the row grain, key, expected ranges, clock, known source delays and intended decision. For receipt handling, inspect duplicate submission IDs, negative totals, impossible review timestamps and whether a later correction overwrote an earlier value. A histogram can expose a spike; it cannot explain whether the spike is a real operational event or an import defect. Ask the source owner before deleting a cluster.

Inspect training data in slices

Use counts and quantiles by submission month and region. For a predictive task, keep the final test partition closed while investigating anomalies and choosing transformations. Repeatedly viewing test outcomes while revising features turns the test into another development set. For a descriptive report, use all permitted rows but freeze an extraction cutoff and record the filters. A reproducible snapshot] makes later comparisons meaningful.

Turn observations into checks

If the exploration finds a burst of duplicate IDs after a provider migration, add a duplicate-rate check with an owner and an alert threshold. If receipt amounts have a long tail, report a median and high quantile alongside the mean. Inspect sample sizes for small slices before making a confident claim. Correlation between queue length and processing time may reflect shared load; it does not prove that one causes the other.

Keep the analysis legible

Save a compact manifest of filters, code revision, extraction time and counts. An analyst should be able to recreate each chart from that manifest. Store aggregate outputs where possible and apply the dataset access policy to row-level extracts. Uncertainty intervals] become useful only after the population and sampling procedure are clear.

Implementation

python
checks = {
    "rows": len(receipts),
    "duplicate_ids": int(receipts["submission_id"].duplicated().sum()),
    "negative_amounts": int(receipts["amount_paid"].lt(0).sum()),
    "late_reviews": int((receipts["reviewed_at"] < receipts["submitted_at"]).sum()),
}
assert checks["duplicate_ids"] == 0
assert checks["negative_amounts"] == 0
print(receipts.groupby("submission_month")["amount_paid"].quantile([0.5, 0.9]))

Performance and operating cost

A single scan for each check costs O(N) time; sorting for exact quantiles may cost O(N log N) and substantial memory. Sample for early investigation, then validate final claims on the full defined population.

Common Mistakes

  • Do not pick transformations by repeatedly reading the sealed test set.
  • Do not delete an outlier only because it looks unusual.
  • Do not report a slice without its row count and extraction cutoff.

Read next

Connected implementation

Continue the workflow: Column profiles and domain constraints before analysis.

data-science
exploratory-analysis-without-peeking
Storage details