Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Pipeline stages: cache by complete input identity

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A cached training stage is reusable only when its code, configuration, data and environment inputs are unchanged.

Define a stage as a contract

A receipt training pipeline may validate data, assemble features, split cohorts, fit a model and evaluate it. Each stage needs declared inputs, output artifacts and failure conditions. A stage that reads a mutable “latest” table outside its declared inputs cannot be replayed from its run record. Put the dataset snapshot, transformation revision, split policy and runtime image digest in the stage identity. Training manifests connect those identities to the resulting model. Scheduling tells the stage when to run; it does not make its outputs reproducible.

Compute cache keys from all effective inputs

Caching a feature stage by source-code hash alone is unsafe when the source partition changes. Include exact data artifact digests, code revision, configuration values and environment identity. Conversely, a display-only run label should not invalidate expensive computation. Record the cache hit and its producer run so an auditor can distinguish a reused artifact from a freshly computed one. Randomized stages need a declared seed; nondeterministic hardware kernels may still prevent byte-for-byte reproduction, which should be stated rather than concealed.

Keep stages atomic and inspectable

Write output to a temporary location, validate it and publish an immutable artifact only after success. A failed stage should not leave a half-written dataset at the path another stage reads. Avoid rerunning the entire pipeline when one evaluation stage fails; reuse valid immutable parents. But if an upstream data correction changes its digest, every dependent stage must be reconsidered. Feature contract changes are particularly easy to omit from cache identity when they live in a separate repository.

Test invalidation before trusting the cache

Build a fixture where only data changes, one where only feature configuration changes and one where only a run label changes. The first two must miss the cache; the last may hit it. Compare artifact digests and stage metadata after each run. The replay project asks for the exact inputs that produced a candidate, plus an explicit explanation when a replay is semantically equal but not byte-identical.

Implementation

python
import hashlib
import json

def stage_cache_key(stage_name, code_digest, input_digests, config, image_digest):
    if not stage_name or not code_digest or not image_digest:
        raise ValueError("stage, code and image identity are required")
    payload = {"stage": stage_name, "code": code_digest,
               "inputs": input_digests, "config": config,
               "image": image_digest}
    encoded = json.dumps(payload, sort_keys=True, separators=(",", ":")).encode()
    return hashlib.sha256(encoded).hexdigest()

inputs = {"receipts": "sha256:data-r47"}
first = stage_cache_key("features", "sha256:code-r8", inputs,
                        {"age_cap_days": 3650}, "sha256:image-r6")
assert first == stage_cache_key("features", "sha256:code-r8", inputs,
                                {"age_cap_days": 3650}, "sha256:image-r6")
assert first != stage_cache_key("features", "sha256:code-r8",
                                {"receipts": "sha256:data-r48"},
                                {"age_cap_days": 3650}, "sha256:image-r6")

Performance and operating cost

Canonicalizing m metadata fields takes O(m) space, while hashing takes O(B) time for B encoded metadata bytes. Computing input digests can require a full scan of large data before a cache lookup. Reuse verified immutable artifact digests when available; a mutable path or timestamp is not a trustworthy substitute.

Common Mistakes

  • Caching on code revision while ignoring changed data.
  • Using a mutable latest-table name as an input identity.
  • Publishing partial stage outputs after failure.
  • Claiming exact reproducibility from a seed while the execution kernel remains nondeterministic.

Read next

Continue the workflow: Training resource gates: bound cost before starting a run.

ai-data
mlops
Storage details