An ML release needs small deterministic checks and a few realistic integration checks, each tied to a different failure boundary.
ML tests: separate code, data, model and service failures
Start with the cheapest contract checks
A receipt model can pass an accuracy test while its service rejects every request because the amount field changed units. Test parsers, feature transforms, schema admission, missing-value behavior and output shape with tiny fixed fixtures before spending time on training. Keep each fixture under a stated contract version and include invalid as well as valid records. Feature admission owns the input boundary; these tests prove a release still implements it.
Add data and training checks
Validate a frozen data snapshot for row identity, label maturity, split isolation, class counts and feature availability. Then train a small deterministic smoke model to verify the pipeline can finish and publish complete artifacts. The smoke run does not establish production quality. Run the full candidate evaluation on a fixed holdout, protected slices and cost limits only after the cheap checks pass. Training replay explains why the data and split identities belong in the test record.
Test the deployed interface
Load the exact candidate artifact in a staging service with the production request schema. Check one accepted request, one rejected request, a deadline breach, a fallback and a rollback load. A model file that loads in a notebook can still fail in a different image or feature environment. Shadow traffic can reveal operational behavior, but it should not be the first time anyone checks the request contract. Latency budgets turn a vague speed expectation into a measurable service test.
Report failure at the right layer
A test report should name the artifact digest, fixture revision, environment image and failing boundary. Do not bury a schema failure under “accuracy regression.” Keep test inputs small enough for every code change, then reserve full training and load tests for changes that can affect them. The release-gate project exercises code, data, model and service failures separately so one broad pass/fail flag cannot hide which layer broke.
Implementation
def release_checks(results):
ordered = ("schema", "transform", "data", "training-smoke",
"quality", "service-smoke", "rollback")
for check in ordered:
if check not in results:
return {"state": "hold", "reason": "missing-" + check}
if not results[check]:
return {"state": "hold", "reason": "failed-" + check}
return {"state": "ready", "checks": len(ordered)}
passing = {name: True for name in ("schema", "transform", "data",
"training-smoke", "quality", "service-smoke", "rollback")}
assert release_checks(passing)["state"] == "ready"
assert release_checks({**passing, "schema": False}) == {
"state": "hold", "reason": "failed-schema"}
Performance and operating cost
Checking k results is O(k) time and O(1) extra space. The expensive checks are full training, evaluation and load testing; ordering cheap failures first reduces wasted compute. A Boolean result is only useful if it is tied to the exact artifact, dataset and environment tested.
Common Mistakes
- Using one accuracy number as the entire release test.
- Running costly training before cheap schema checks.
- Testing a notebook loader but not the deployed request interface.
- Reusing a passing test report for a different artifact digest.
Read next
- Environment promotion: keep the tested artifact and contract together
- Project: gate a receipt model through CI and staging
- Feature contracts: admit only usable inference records
- Training replay: freeze the cohort, split and runtime
- Model promotion: require evidence before changing the serving pointer
Continue the workflow: Migrate inference clients with compatibility tests and usage evidence.
