A cache needs tests for each identity that can change a decision, including a silent feature update and a threshold-only release.
Cache invalidation tests for model, feature and policy changes
List every invalidation trigger
A new model digest, feature revision, preprocessor, tenant permission or policy threshold can change the answer. Build a table of which values are part of a score cache versus a final-decision cache. If only scores are cached, reapply current policy and log that policy revision on every hit. If final decisions are cached, policy revision belongs in the key. Identity-rich keys make most releases naturally miss old entries without global deletion.
Test time as an input
An entry can be correct when written and wrong after a feature budget expires. Validate the entry expiry against the current clock and any embedded feature watermark; do not simply reset its TTL on every hit. A stale entry must trigger recomputation or the documented fallback. Freshness limits apply even if the model artifact is unchanged. A client retry immediately after a timeout can use an idempotency record for the same logical decision; that is a separate promise from a general prediction cache.
Probe cross-region and partial invalidation
Warm the same request in two replicas, change the feature revision and confirm both reject the old key. Then fail over to a region whose cache has an earlier entry. The fallback region must check model and feature identity rather than trust the cache presence. A broadcast invalidation can be delayed or lost, which is why versioned keys and short expiry are useful. Failover identity should be checked before routing cached responses to customers.
Audit actual hits
Record entry creation and expiry, model digest, feature revision, policy revision and hit or recompute decision in restricted events. Count stale rejects and conflicting recomputations. A high hit rate is a cost metric, not a correctness metric; a sampled fresh recompute should compare routes and investigate any divergence. The applied project forces a merchant feature change and a policy-only release, then proves the original route cannot be reused.
Implementation
def reusable_cached_decision(entry, current, now_seconds):
if now_seconds >= entry["expires_at"]:
return "miss:expired"
required = ("tenant", "feature_revision", "model_digest",
"policy_revision", "request_digest")
if any(entry.get(field) != current.get(field) for field in required):
return "miss:identity"
return "hit"
entry = {"tenant": "merchant-47", "feature_revision": "f3",
"model_digest": "risk-47", "policy_revision": "p4",
"request_digest": "req-82", "expires_at": 1082}
current = {key: value for key, value in entry.items() if key != "expires_at"}
assert reusable_cached_decision(entry, current, 1047) == "hit"
assert reusable_cached_decision(entry, {**current,
"policy_revision": "p5"}, 1047) == "miss:identity"
Performance and operating cost
The gate compares a fixed set of fields in O(f) time and O(1) extra space for f fields. Versioned cache keys increase entry count during rollouts until old entries expire, while very short TTLs lower hit rate. A fresh recompute sample costs additional model work but can reveal silent identity bugs that hit-rate metrics cannot.
Common Mistakes
- Refreshing TTL on every hit until old feature state lives forever.
- Clearing one replica while another continues serving old entries.
- Using hit rate as evidence of decision correctness.
- Confusing an idempotent retry record with a reusable prediction cache.
Read next
- Prediction caches: key by every input that changes the decision
- Project: prove receipt prediction-cache identity across releases
- Online feature freshness: use event and availability clocks
- Inference retries: bound repeated work and preserve one decision
- Region failover for inference: match model and feature state
