A cached decision is valid only for the same authorized input, model, feature state and policy under a defined freshness window.
Prediction caches: key by every input that changes the decision
Decide whether caching is valid
A receipt may be rescored after a merchant update or a policy change. Caching can save compute for repeated identical requests, but a cache hit may hide new evidence. Before enabling it, define which inputs are immutable, which features can change and the maximum age acceptable for the decision. Feature freshness sets an upper bound; a cache lifetime cannot make stale features fresh. Distinguish a reusable model score from a final route, because a policy threshold may change without retraining.
Construct an identity-rich key
Scope a key to tenant and authorization context, canonical request digest, feature snapshot or watermark, preprocessing revision, model digest and policy revision when caching the final route. A cached score can omit policy revision only if policy is reapplied on every hit and its response is logged with the current policy. Never share a result across tenants because two receipts happen to contain the same bytes. The API boundary names the response semantics; artifact identity pins model bytes.
Set a defensible expiry
Use the shortest relevant lifetime among feature freshness, business decision validity and data retention. A cache entry should carry creation time, expiry and the identities it covers; reject it when any required identity changes. A global “clear cache” during a release is not enough if some replicas or regions miss the invalidation. Versioned keys make old entries unreachable without relying on perfect broadcast. Region failover needs the same digest and watermark checks before reusing a cached route.
Measure hit quality, not just hit rate
Track hits, misses, stale rejections, identity mismatches and recompute latency by bounded model and route dimensions. Sample a limited set of hits for fresh recomputation to detect an invalid assumption, without turning routine sampling into customer-visible decisions. Invalidation tests exercise policy and feature changes; the project catches a high hit rate that conceals old merchant data.
Implementation
from hashlib import sha256
import json
def decision_cache_key(tenant, request_bytes, feature_revision,
model_digest, policy_revision):
fields = (tenant, sha256(request_bytes).hexdigest(), feature_revision,
model_digest, policy_revision)
canonical = json.dumps(fields, separators=(",", ":"), ensure_ascii=False)
return sha256(canonical.encode()).hexdigest()
base = decision_cache_key("merchant-47", b"receipt-82", "features-r3",
"risk-47", "policy-4")
assert base == decision_cache_key("merchant-47", b"receipt-82",
"features-r3", "risk-47", "policy-4")
assert base != decision_cache_key("merchant-47", b"receipt-82",
"features-r4", "risk-47", "policy-4")
Performance and operating cost
Hashing b request bytes and v identity bytes costs O(b + v) time and O(v) temporary space for the canonical identity. Cache storage scales with distinct live keys and response size. JSON array encoding preserves field boundaries; a raw delimiter join could make different identity tuples collide.
Common Mistakes
- Keying only on raw request bytes while feature state changes.
- Reusing a final route after the threshold policy changes.
- Sharing cache entries across tenants or authorization scopes.
- Using cache TTL longer than the shortest feature freshness budget.
Read next
- Cache invalidation tests for model, feature and policy changes
- Project: prove receipt prediction-cache identity across releases
- Online feature freshness: use event and availability clocks
- Score thresholds are release policy, not model metadata
- Region failover for inference: match model and feature state
