An expand-contract migration changes a shared data contract in stages so old and new readers can coexist until measured evidence permits the old field to retire.
Expand-contract schema migration
Start with the reader inventory
A raw event carries amount_cents, while a new producer wants amount_minor_units and currency. The rename is not merely a DDL command: dashboards, streaming jobs, backfills and exports may depend on the old name. Inventory readers by owner, query identity and last observed use. Mark unknown readers as a release blocker until access logs or owner review resolve them. Consumer rollout supplies the compatibility plan.
Expand before changing writers
Add the new nullable fields and publish a schema version that accepts both representations. Write both fields from one validated amount in the producer; a second independent calculation can produce disagreement. Test the serialized event and persisted table, since compatibility in one layer does not guarantee compatibility in the other. Keep the old field available while lagging readers and retained historical objects still require it.
Backfill with a fence
A large table should be repaired in bounded partitions rather than one unbounded rewrite. Pin input snapshot and transform revision, write candidates to new generations, and record a high-water position. Live updates can race a backfill: compare source positions or use a change log so a stale repair never overwrites a later value. The publication boundary prevents a half-migrated partition from reaching readers.
Move readers using evidence
Switch one consumer at a time to the new field, then compare counts, null rates and monetary totals by event date and currency. If historical rows cannot be converted, surface an explicit unknown state; do not invent a zero amount. Observe reader lag through the maximum replay horizon before claiming the old field is unused. A rollback should restore the previous reader pointer without losing newly written events.
Contract only after the horizon
Stop dual-writing after every required reader has moved and historical replay tests pass. Remove the old field in a later release, not in the same deployment that switches readers. Rehearse a delayed reader, a backfill retry and an emergency rollback. Attach the schema IDs, reconciliation output and approval to the release manifest so a future operator can reconstruct the decision.
Implementation
events = [
{"id": "pay-47", "amount_cents": 2375, "currency": "USD"},
{"id": "pay-48", "amount_cents": 6420, "currency": "USD"},
]
def expand_payment(event):
migrated = dict(event)
migrated["amount_minor_units"] = event["amount_cents"]
return migrated
expanded = [expand_payment(event) for event in events]
assert sum(row["amount_minor_units"] for row in expanded) == 8795
assert all(row["amount_cents"] == row["amount_minor_units"] for row in expanded)Performance and operating cost
The reference pass is O(N) time and O(N) output space for N events. At scale, dual fields increase storage and transport size during the overlap, while a historical rewrite consumes I/O proportional to retained bytes. Partitioned checkpoints limit retry work; they do not remove the need to reconcile concurrent updates and old reader behavior.
Common Mistakes
- Do not drop a field because the newest application build no longer reads it.
- Do not let a stale backfill replace a newer live update.
- Do not equate a nullable new column with a complete historical migration.
