Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Backfill prompts: compare count and field-level parity

Last updated: 5 Oct 202611 min read
tutorial
AdvancedBy AITrove Editorial

Count parity is necessary but weak. A validation prompt should request the fixed source population, mapped and quarantined counts, destination unique-key count, duplicate-key check, field comparison, and mismatch ledger. Hashes help only when both sides use the same canonical encoding and field set. Keep row IDs for every mismatch. Reconcile all groups to the source total; a target with the right number of rows can still contain wrong statuses.

Operational case

Harbor has 47,000 source rows: 46,940 mapped and 60 quarantined. The target has 46,940 rows, so eligible count parity passes. A keyed status comparison finds three wrong values, leaving 46,937 matching. The model's report states both facts and withholds the cutover claim. An engineer checks whether mapping logic or retry ordering caused the three errors, then reruns the affected and full-population checks.

Output
47,000 = 46,940 mapped + 60 quarantined
46,940 target = 46,937 matching + 3 mismatched
Count gate: pass | field-value gate: fail
Cutover: blocked

Performance and review cost

Sorting both N-row sets by key costs O(N log N); an indexed merge is O(N) after index preparation. Hash comparison is O(N) expected time and O(N) memory. The extra read pass reveals errors that a total cannot, and the three keyed IDs make the repair check specific.

Common Mistakes

  • Do not equate equal counts with equal values.
  • Do not compare hashes with different field encodings.
  • Do not omit quarantined rows from source reconciliation.

Connected lessons

prompt engineering
database migration
Storage details