A quality gate checks both individual records and batch-level invariants, retaining rejected rows for diagnosis rather than silently shrinking the dataset.
Data quality gates: quarantine bad rows and reconcile complete batches
Use two levels of checks
A row rule rejects a missing receipt ID, invalid currency or negative amount. A batch rule checks source count, accepted count, rejected count, unique keys and expected partitions. A file can have valid rows but still be incomplete because the transfer stopped halfway. Therefore a row parser alone is not a release gate. The source contract] defines which failures are fatal and which may be quarantined.
Keep the rejected record recoverable
Store a bounded copy of the raw record, source position, rule version and reason in an access-controlled quarantine. Redact sensitive fields from alert messages. If the producer fixes the record, replay it with its original identity; do not assign a new identity that could double count it. Expire quarantine data according to retention policy, but first preserve aggregate failure counts for trend analysis.
Reconcile before publish
Assert source_rows = accepted_rows + rejected_rows + intentionally_skipped_rows. Count distinct keys and compare totals that should be conserved, such as amount in cents, only when transformation semantics permit it. Pause publication when a required partition is missing or rejection rate crosses an agreed threshold. A warning alone can allow an incomplete dashboard to look authoritative.
Run failure drills
Test a truncated input, duplicate key, malformed timestamp and a whole missing partition. Verify which cases stop publication and which go to quarantine. Check that a retry does not duplicate accepted rows. Idempotent loads] keep repair from creating a second failure.
Implementation
def reconcile_batch(source_count, accepted_count, rejected_count, skipped_count):
accounted = accepted_count + rejected_count + skipped_count
if accounted != source_count:
raise ValueError(f"unaccounted records: {source_count - accounted}")
if source_count and rejected_count / source_count > 0.02:
raise ValueError("receipt rejection budget exceeded")
return {"source": source_count, "accepted": accepted_count,
"rejected": rejected_count, "skipped": skipped_count}Performance and operating cost
Row checks are O(N); distinct-key checks use O(K) memory for K keys in a local hash set or an indexed database. Quarantine storage and review effort grow with failure volume.
Common Mistakes
- Do not assume valid rows imply a complete file.
- Do not discard invalid rows before recording why they failed.
- Do not publish a partially reconciled batch as a normal update.
Read next
- Data source contracts: preserve raw records before transformation
- Idempotent loads: commit target rows and extraction progress together
- Backfills: rebuild history without exposing a half-written result
- Missing data policy: distinguish absence from a measured zero
Continue the workflow: Spatial joins: polygons, boundaries and duplicate matches.
Continue the workflow: Synthetic relational data: keys, chronology and integrity.
Continue the workflow: Outlier investigation: impossible values, rare events and quarantines.
Continue the workflow: SQL reconciliation: prove the cohort survived each transformation.
Continue the workflow: Dependency readiness and source freshness.
Continue the workflow: Semantic metric contracts and reconciliation.
Continue the workflow: Cohort baselines and denominator drift.
Continue the workflow: Migration reconciliation and cutover gates.
Continue the workflow: Cross-engine shadow reads and discrepancy triage.
Continue the workflow: Spatial candidate and exact-predicate joins.
