Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Prediction caches: key by every input that changes the decision

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A cached decision is valid only for the same authorized input, model, feature state and policy under a defined freshness window.

Decide whether caching is valid

A receipt may be rescored after a merchant update or a policy change. Caching can save compute for repeated identical requests, but a cache hit may hide new evidence. Before enabling it, define which inputs are immutable, which features can change and the maximum age acceptable for the decision. Feature freshness sets an upper bound; a cache lifetime cannot make stale features fresh. Distinguish a reusable model score from a final route, because a policy threshold may change without retraining.

Construct an identity-rich key

Scope a key to tenant and authorization context, canonical request digest, feature snapshot or watermark, preprocessing revision, model digest and policy revision when caching the final route. A cached score can omit policy revision only if policy is reapplied on every hit and its response is logged with the current policy. Never share a result across tenants because two receipts happen to contain the same bytes. The API boundary names the response semantics; artifact identity pins model bytes.

Set a defensible expiry

Use the shortest relevant lifetime among feature freshness, business decision validity and data retention. A cache entry should carry creation time, expiry and the identities it covers; reject it when any required identity changes. A global “clear cache” during a release is not enough if some replicas or regions miss the invalidation. Versioned keys make old entries unreachable without relying on perfect broadcast. Region failover needs the same digest and watermark checks before reusing a cached route.

Measure hit quality, not just hit rate

Track hits, misses, stale rejections, identity mismatches and recompute latency by bounded model and route dimensions. Sample a limited set of hits for fresh recomputation to detect an invalid assumption, without turning routine sampling into customer-visible decisions. Invalidation tests exercise policy and feature changes; the project catches a high hit rate that conceals old merchant data.

Implementation

python
from hashlib import sha256
import json

def decision_cache_key(tenant, request_bytes, feature_revision,
                       model_digest, policy_revision):
    fields = (tenant, sha256(request_bytes).hexdigest(), feature_revision,
              model_digest, policy_revision)
    canonical = json.dumps(fields, separators=(",", ":"), ensure_ascii=False)
    return sha256(canonical.encode()).hexdigest()

base = decision_cache_key("merchant-47", b"receipt-82", "features-r3",
                          "risk-47", "policy-4")
assert base == decision_cache_key("merchant-47", b"receipt-82",
                                  "features-r3", "risk-47", "policy-4")
assert base != decision_cache_key("merchant-47", b"receipt-82",
                                  "features-r4", "risk-47", "policy-4")

Performance and operating cost

Hashing b request bytes and v identity bytes costs O(b + v) time and O(v) temporary space for the canonical identity. Cache storage scales with distinct live keys and response size. JSON array encoding preserves field boundaries; a raw delimiter join could make different identity tuples collide.

Common Mistakes

  • Keying only on raw request bytes while feature state changes.
  • Reusing a final route after the threshold policy changes.
  • Sharing cache entries across tenants or authorization scopes.
  • Using cache TTL longer than the shortest feature freshness budget.

Read next

ai-data
mlops
Storage details