A backfill computes corrected partitions under a versioned rule and publishes them only after quality checks pass as a complete replacement.
Backfills: rebuild history without exposing a half-written result
Separate build from visibility
A new receipt-tax rule requires rebuilding six months of aggregates. Updating the visible table one month at a time makes dashboards mix old and new rules. Build corrected partitions in a staging namespace, validate totals and key coverage, then switch a version pointer or replace partitions through the warehouse’s atomic publish mechanism. Readers should see either the old complete version or the new complete version.
Pin inputs and rules
Record source snapshots, code revision, transform parameters and affected partition range. If the source receives late changes while the backfill runs, choose a consistent cutoff and apply subsequent changes after publication. Running a backfill against moving data without a cutoff creates a result no one can reproduce. Snapshot manifests] document that choice.
Budget replay pressure
A backfill can compete with live ingest for CPU, I/O and database locks. Limit concurrency and measure lag in the live pipeline. Maintain a rollback pointer to the prior version, but do not assume rollback can undo irreversible downstream exports. List consumers and cache invalidation before the switch.
Verify the repaired history
Compare row counts, key uniqueness, rejection counts and representative totals against the old version. Recompute at least one downstream report and record the expected change. Deliberately fail a staged partition and confirm that no pointer swap occurs. Quality gates] make publication conditional on complete evidence.
Implementation
backfill_manifest = {
"version": "receipt-rule-v47",
"source_snapshot": source_snapshot_id,
"partitions": list(affected_months),
"validated": all(partition_checks.values()),
}
if not backfill_manifest["validated"]:
raise RuntimeError("backfill failed quality gates")
publish_version_pointer(backfill_manifest["version"])Performance and operating cost
A full backfill reads and rewrites O(H) historical rows for H affected records, plus temporary storage for a parallel version. An atomic pointer switch is cheap; the expensive part is the validated rebuild and downstream reconciliation.
Common Mistakes
- Do not expose a mixture of old and new transformation rules.
- Do not rebuild against an unpinned moving source without a cutoff policy.
- Do not publish when one required partition failed validation.
Read next
- Data quality gates: quarantine bad rows and reconcile complete batches
- Idempotent loads: commit target rows and extraction progress together
- Event time and late arrivals: close windows with an explicit correction policy
- Reproducible analysis snapshots: pin data, code and cutoff together
Continue the workflow: Index freshness: publish complete revisions and remove stale chunks.
Continue the workflow: Dashboard filters and provenance: make every view reproducible.
Continue the workflow: Feature operations: backfills, deletion and atomic publication.
Continue the workflow: Replay manifests and audit trails.
