A feature rename or unit change is a release with consumers, not an isolated edit to a training column.
Feature schema evolution: keep producers and rollback models compatible
Inventory every consumer of a field
A receipt amount feeds online scoring, batch replay, dashboards and perhaps the previous model retained for rollback. Before changing `amount_minor`, list each consumer and the contract version it accepts. Adding an optional field may be safe for a tolerant reader, yet removing a field can break the rollback path. Run lineage records the training input version; runtime compatibility must still be tested against live producers.
Distinguish additive changes from changed meaning
Changing an integer amount from cents to thousands of cents without renaming it preserves the type and destroys the meaning. Introduce a new field and migrate consumers in stages. Keep both fields only while both are derived from one authoritative source and checked for equivalence. Document when the old field stops being emitted. A name alias can make a deployment pass while leaving historical replay wrong. Admission rules should reject an unexpected contract version rather than guess units.
Prove temporal consistency
Training rows need values available at their prediction time. An updated feature view may backfill corrected values that did not exist when older decisions were made. Store event time and availability time, and define whether replay means historical truth or information available then. Point-in-time joins address the latter. A schema migration test should include old and new event windows, late-arriving updates and the previous serving model.
Release in a reversible order
First deploy readers that understand the new field while still accepting old records. Then emit both fields, compare values, move model candidates to the new contract and finally retire the old field after the rollback window closes. Gate each step on mismatch and unknown-version rates. The receipt admission project tests a unit migration with a deliberately stale consumer so the hold decision is visible.
Implementation
def compatible_contract(producer, consumers):
failures = []
for consumer_name, required in consumers.items():
for field, unit in required.items():
if producer.get(field) != unit:
failures.append((consumer_name, field, unit))
return {"ready": not failures, "failures": failures}
producer = {"amount_minor": "cent", "merchant_age_days": "day",
"amount_thousands": "thousand-cent"}
consumers = {"candidate-r8": {"amount_thousands": "thousand-cent"},
"rollback-r7": {"amount_minor": "cent"}}
assert compatible_contract(producer, consumers)["ready"]
assert not compatible_contract({"amount_thousands": "thousand-cent"},
consumers)["ready"]
Performance and operating cost
Checking c consumers against their required fields costs O(total required fields) time and O(failures) space. Dual-writing two representations adds producer work and temporary storage; it also creates a comparison obligation. The example checks declared fields and units, while a real release must also test value equivalence and historical availability.
Common Mistakes
- Changing a unit without changing the field identity.
- Deleting the old field while a rollback model still requires it.
- Backfilling future-known values into a historical training replay.
- Keeping dual writes forever without a retirement condition.
Read next
- Feature contracts: admit only usable inference records
- Project: release a versioned receipt feature admission gate
- Training manifests: link data, code, configuration and artifact
- Training-serving parity: compare feature values at one prediction clock
- Historical feature retrieval: point-in-time joins without future leakage
Continue the workflow: Inference API contracts: version the decision, not only the payload.
