Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Project: replay a receipt model from frozen training evidence

Last updated: 7 Oct 20265 min read
project
AdvancedBy AITrove Editorial

Build a small training pipeline whose artifacts can be explained, invalidated and compared across reruns.

Define the training contract

The pipeline consumes a frozen receipt snapshot, an outcome revision cutoff and a feature-contract version. It validates inputs, builds features, assigns stable merchant-level splits, trains a model and evaluates a fixed holdout. Each stage publishes an immutable artifact only after validation passes. Stage identities include data and configuration, while replay evidence records the eligible population and runtime.

Construct invalidation fixtures

Run the pipeline twice with identical inputs; the second run may reuse verified artifacts. Then change one label revision, one feature threshold and one presentation-only run name. Require the first two to invalidate dependent stages and the last to leave computation unchanged. Include a failed training attempt that writes half an artifact; the publishing step must not expose it. Record cache hits and their producer run IDs so reuse is visible rather than mistaken for fresh work.

Compare candidate and incumbent fairly

Evaluate the new artifact and current production model on the same frozen holdout and feature availability rules. Check overall loss, urgent-case recall and operational latency. A better average metric cannot waive a protected slice limit. Keep test feedback out of repeated parameter tuning. Promotion gates consume the manifest and evaluation results for the exact new artifact digest.

Deliver the replay report

Produce a manifest with input digests, stage versions, split membership digest, environment digest, model digest and metric-code revision. Report whether a rerun matched bytes, matched metrics within a declared tolerance or differed. Explain the earliest changed stage. A reviewer should be able to reproduce the candidate without consulting a mutable “latest” table or guessing which dependency image ran.

Implementation

python
def replay_comparison(expected, observed, metric_tolerance=0.002):
    identity_fields = ("data_digest", "split_digest", "code_digest",
                       "image_digest", "feature_contract")
    changed = [field for field in identity_fields
               if expected.get(field) != observed.get(field)]
    if changed:
        return {"state": "different-inputs", "changed": changed}
    if expected["model_digest"] == observed["model_digest"]:
        return {"state": "byte-identical"}
    delta = abs(expected["holdout_loss"] - observed["holdout_loss"])
    return {"state": "metric-equivalent" if delta <= metric_tolerance
            else "different-result", "loss_delta": delta}

base = {"data_digest": "d47", "split_digest": "s82", "code_digest": "c5",
        "image_digest": "i6", "feature_contract": "receipt-v3",
        "model_digest": "m8", "holdout_loss": 0.173}
assert replay_comparison(base, dict(base))["state"] == "byte-identical"
assert replay_comparison(base, {**base, "data_digest": "d48"})[
    "state"] == "different-inputs"

Performance and operating cost

Comparing a fixed-size manifest is O(f) time and O(f) space for f identity fields. Training and evaluation dominate runtime; retaining immutable snapshots and artifacts adds storage cost. Metric equivalence is a limited claim, not proof that the same decision function was produced for every input.

Common Mistakes

  • Calling a cached stage freshly executed.
  • Publishing incomplete output after a failed stage.
  • Comparing models on different holdout populations.
  • Using a matching score to claim identical artifact bytes.

Read next

Continue the workflow: Project: resume receipt-model training under a compute budget.

ai-data
mlops
Storage details