Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Schema compatibility and consumer rollout

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A schema contract defines how producers encode records and which historical and future records each consumer can read.

Separate wire compatibility from meaning

Adding a nullable field can be harmless for a decoder yet still break a downstream calculation that assumes the field is always present. Renaming an amount field while keeping its numeric type may preserve bytes but change business meaning. Record field identity, units, null behavior, keys, and deletion meaning in the contract. Raw landing should retain the original payload and its schema version so a corrected parser can replay it.

Choose the reader and writer order

Backward compatibility asks whether a new reader can consume older records; forward compatibility asks whether an older reader can consume records from a new writer. A rollout with persistent historical files needs a reader that understands old versions. A mixed fleet of consumers may also need old readers to understand new writes. Check the chosen compatibility rule against every supported version when old files or clients can survive more than one release.

Exercise an actual consumer fixture

A registry check cannot test that a billing report still interprets cents as cents. Keep representative payloads from each supported version, including nulls, deleted entities, and unexpected enum values. Decode them with the candidate consumer, validate the resulting normalized record, then run the aggregate that matters. A failing contract test should block publication before quality gates see the changed data.

Publish in stages

First deploy consumers that accept both old and new payloads. Then permit producers to emit the new version and monitor decode errors and unknown-field counts. Only retire old parsing after the retention horizon and dependent consumers have moved. A breaking change belongs in a new stream or versioned dataset with an explicit migration, not an unannounced reinterpretation of the existing one.

Make failure reversible

Keep the last known-good schema and a rollback switch for the writer. Rolling back the binary alone does not remove records already emitted in the new format. If a new writer emitted 47 incompatible records, quarantine those records by event ID, repair or translate them, and replay from the raw log with the same deduplication key.

Implementation

python
SUPPORTED = {2: {"order_id", "amount_cents"},
             3: {"order_id", "amount_cents", "currency"}}

def normalize_order(payload, schema_version):
    if schema_version not in SUPPORTED:
        raise ValueError("unsupported order schema")
    if not SUPPORTED[schema_version].issubset(payload):
        raise ValueError("missing required field")
    if not isinstance(payload["amount_cents"], int):
        raise TypeError("amount must be integer cents")
    return {"order_id": payload["order_id"],
            "amount_cents": payload["amount_cents"],
            "currency": payload.get("currency", "INR")}

assert normalize_order({"order_id": "ord-47", "amount_cents": 7350}, 2)["currency"] == "INR"

Performance and operating cost

A fixture suite is O(V × R) over V supported schema versions and R representative records. Real cost comes from keeping old readers and migration paths until retention ends. A registry compatibility result is necessary for encoded records, but semantic checks and staged rollout still need their own tests.

Common Mistakes

  • Do not assume an optional wire field is optional in every consumer calculation.
  • Do not equate an accepted schema change with a safe unit or enum change.
  • Do not remove a reader for old data while retained files still use it.

Read next

Continue the workflow: Conformed dimensions and surrogate keys.

Continue the workflow: Dataset ownership and change approval.

Continue the workflow: Schema registry and transitive compatibility.

Continue the workflow: Expand-contract schema migration.

Continue the workflow: Cross-engine SQL semantic differences.

Continue the workflow: Direct and indirect column lineage.

ai-data
data-engineering
Storage details