Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Region failover for inference: match model and feature state

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A healthy standby endpoint is useful only when its model, policy and feature freshness are suitable for the decisions it will serve.

Treat the region as a complete decision path

An inference request may cross a gateway, feature store, model worker, policy service and decision log. Moving only the model endpoint to a second region can produce a response with stale features or an older threshold. Maintain a manifest for each region: approved artifact digest, API contract, feature snapshot or watermark, policy revision, log sink and readiness state. Artifact lineage identifies the model; feature parity tests whether inputs carry the intended meaning across environments.

Decide what lag is acceptable

The standby may have a replicated feature stream that trails the primary by minutes. Define a recovery point for feature state and a separate maximum age for individual features. A failover gate must compare the standby watermark to request time and the feature-specific freshness budget. Some decisions may route to manual review when lagged; others may be rejected. Do not present a successful score from stale state as an ordinary decision. Freshness budgets and fallback rules need to agree.

Pin the served identity

Record the region, artifact digest, policy revision, feature watermark and decision ID for every served response. A mutable alias such as production can point to different bytes in each region during deployment, so failover should compare resolved digests rather than alias text. Promotion must replicate and verify the package before traffic moves. If the standby has the previous approved digest, state whether that is an authorized rollback or a hold. A reversible pointer protects routing, but it cannot make a missing artifact appear.

Test the unhappy path

Run a planned switch when the standby is current, then simulate an absent artifact and a lagged feature watermark. The first switch should succeed; the others should hold or use a named safe route. Keep the primary isolated long enough to check whether it continued issuing decisions during failover. The failover project tests duplicate decision IDs and recovery records. A load balancer health check that returns 200 is weaker than a scored canary with verified identity and freshness.

Implementation

python
def standby_route(primary, standby, maximum_lag_seconds=47):
    if primary["api_contract"] != standby["api_contract"]:
        return "hold:api-contract"
    if primary["approved_digest"] != standby["approved_digest"]:
        return "hold:model-digest"
    lag = primary["feature_watermark"] - standby["feature_watermark"]
    if lag < 0 or lag > maximum_lag_seconds:
        return "manual-review:feature-lag"
    return "serve:standby"

primary = {"api_contract": "receipt-v2", "approved_digest": "model-47",
           "feature_watermark": 1000}
standby = {**primary, "feature_watermark": 970}
assert standby_route(primary, standby) == "serve:standby"
assert standby_route(primary, {**standby, "approved_digest": "model-46"}) == "hold:model-digest"

Performance and operating cost

The manifest comparison is O(1) time and space. Real failover costs duplicate capacity, artifact replication, fresh feature delivery and periodic drills. A 47-second global lag threshold is illustrative; the correct budget depends on the feature and decision. A healthy endpoint does not guarantee a safe feature watermark or complete decision logging.

Common Mistakes

  • Failing over based only on endpoint health.
  • Comparing alias names without resolving artifact digests.
  • Using a single freshness threshold for every feature.
  • Omitting region identity from decision records.

Read next

Continue the workflow: Cache invalidation tests for model, feature and policy changes.

Continue the workflow: Project: restore a receipt-model registry without changing live decisions.

ai-data
mlops
Storage details