Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Data quality gates: quarantine bad rows and reconcile complete batches

Last updated: 6 Oct 20265 min read
tutorial
IntermediateBy AITrove Editorial

A quality gate checks both individual records and batch-level invariants, retaining rejected rows for diagnosis rather than silently shrinking the dataset.

Use two levels of checks

A row rule rejects a missing receipt ID, invalid currency or negative amount. A batch rule checks source count, accepted count, rejected count, unique keys and expected partitions. A file can have valid rows but still be incomplete because the transfer stopped halfway. Therefore a row parser alone is not a release gate. The source contract] defines which failures are fatal and which may be quarantined.

Keep the rejected record recoverable

Store a bounded copy of the raw record, source position, rule version and reason in an access-controlled quarantine. Redact sensitive fields from alert messages. If the producer fixes the record, replay it with its original identity; do not assign a new identity that could double count it. Expire quarantine data according to retention policy, but first preserve aggregate failure counts for trend analysis.

Reconcile before publish

Assert source_rows = accepted_rows + rejected_rows + intentionally_skipped_rows. Count distinct keys and compare totals that should be conserved, such as amount in cents, only when transformation semantics permit it. Pause publication when a required partition is missing or rejection rate crosses an agreed threshold. A warning alone can allow an incomplete dashboard to look authoritative.

Run failure drills

Test a truncated input, duplicate key, malformed timestamp and a whole missing partition. Verify which cases stop publication and which go to quarantine. Check that a retry does not duplicate accepted rows. Idempotent loads] keep repair from creating a second failure.

Implementation

python
def reconcile_batch(source_count, accepted_count, rejected_count, skipped_count):
    accounted = accepted_count + rejected_count + skipped_count
    if accounted != source_count:
        raise ValueError(f"unaccounted records: {source_count - accounted}")
    if source_count and rejected_count / source_count > 0.02:
        raise ValueError("receipt rejection budget exceeded")
    return {"source": source_count, "accepted": accepted_count,
            "rejected": rejected_count, "skipped": skipped_count}

Performance and operating cost

Row checks are O(N); distinct-key checks use O(K) memory for K keys in a local hash set or an indexed database. Quarantine storage and review effort grow with failure volume.

Common Mistakes

  • Do not assume valid rows imply a complete file.
  • Do not discard invalid rows before recording why they failed.
  • Do not publish a partially reconciled batch as a normal update.

Read next

Continue the workflow: Spatial joins: polygons, boundaries and duplicate matches.

Continue the workflow: Synthetic relational data: keys, chronology and integrity.

Continue the workflow: Outlier investigation: impossible values, rare events and quarantines.

Continue the workflow: SQL reconciliation: prove the cohort survived each transformation.

Continue the workflow: Dependency readiness and source freshness.

Continue the workflow: Semantic metric contracts and reconciliation.

Continue the workflow: Cohort baselines and denominator drift.

Continue the workflow: Migration reconciliation and cutover gates.

Continue the workflow: Cross-engine shadow reads and discrepancy triage.

Continue the workflow: Spatial candidate and exact-predicate joins.

ai-data
data-engineering
Storage details