Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Feature operations: backfills, deletion and atomic publication

Last updated: 5 Oct 20265 min read
tutorial
IntermediateBy AITrove Editorial

Correcting historical features and removing entities require controlled publication so training and serving do not read half-complete state.

Recompute in isolation

Build a corrected feature snapshot into a separate version, preserving the old snapshot for comparison. Reconcile entity counts, missing rates, value distributions and several decision-time examples before switching readers. A backfill can repair a historical data error, but it must not make that repaired value appear available before it was actually known in a past online decision. The publication boundary separates computation from visibility.

Handle deletion explicitly

If an entity is withdrawn or data must be removed, delete or invalidate its online key, cached value and future historical retrieval under the applicable policy. A deletion tombstone must outrank a delayed materialization job; otherwise replay can resurrect the entity. Track downstream training datasets and model artifacts that may require separate retention or remediation decisions.

Plan rollback

Version the source offsets, transform, feature view and online materialization together. If a new snapshot fails parity or freshness checks, return to the last approved snapshot while keeping deletion gates active. A rollback cannot revive a record that must remain withdrawn. Parity evidence belongs in the switch decision.

Exercise a race

Start building version 48 while version 47 serves. Delete merchant M, then let an old worker try to publish M from its stale batch. The serving layer must retain the tombstone. Add a correction to another merchant and confirm the new snapshot appears only after validation. A request trace should identify one coherent snapshot version.

Implementation

python
def publishable_entities(candidate_values, deleted_entity_ids):
    deleted = set(deleted_entity_ids)
    return {entity_id: value for entity_id, value in candidate_values.items()
            if entity_id not in deleted}

Performance and operating cost

Filtering E entities is O(E + D) expected time and O(E + D) memory for D tombstones. Durable systems need transactional or versioned publication plus deletion precedence, not merely this in-memory filter.

Common Mistakes

  • Do not publish partially reconciled backfill output.
  • Do not allow a delayed worker to resurrect a deleted entity.
  • Do not roll back a privacy deletion with an older feature snapshot.

Read next

Continue the workflow: Privacy operations: deletion lineage, access gates and release review.

ai-data
feature-stores
Storage details