Build a release record that carries one model digest from fixture tests through staging smoke checks and rollback rehearsal.
Project: gate a receipt model through CI and staging
Define the candidate package
A receipt-risk candidate has a model digest, feature-contract version, frozen evaluation report and serving image digest. CI runs schema, transform, data, training-smoke and quality checks against that package. The target staging service then runs accepted, rejected and fallback requests. The ML test layers identify which failure boundary a result belongs to instead of flattening every failure into an accuracy result.
Build adversarial cases
Change the amount unit while leaving its type intact, remove one serving permission, swap the artifact after the quality report and make the rollback model incompatible with the new feature producer. Freeze expected hold reasons. A passing average score must not override any of these failures. Store each test result with the package digest and target environment so the report cannot be borrowed by another candidate.
Move through staging deliberately
Verify package bytes, resolve both candidate and rollback digests, run target-environment smoke checks and compare the registry pointer immediately before transition. A concurrent release should halt the operation. If all gates pass, route a small shadow load and then a canary with known limits. Environment promotion keeps the tested bytes and contract together; canary checks cover live operation.
Produce the reviewable result
Report the failed boundary, fixture revision, artifact digest, image digest, data snapshot, target smoke outcome and rollback load result. The project should demonstrate a safe hold for each adversarial case and a clean transition for the complete package. A final “ready” state authorizes the next release step, not an automatic claim that customers saw better outcomes; those require observation and mature labels.
Implementation
def staged_release(ci, staging, candidate_digest, tested_digest):
if candidate_digest != tested_digest:
return {"state": "hold", "reason": "candidate-changed"}
required_ci = {"schema", "transform", "data", "quality"}
required_stage = {"accept", "reject", "fallback", "rollback-load"}
if not required_ci <= {name for name, passed in ci.items() if passed}:
return {"state": "hold", "reason": "ci"}
if not required_stage <= {name for name, passed in staging.items() if passed}:
return {"state": "hold", "reason": "staging"}
return {"state": "ready", "digest": candidate_digest}
ci = {name: True for name in ("schema", "transform", "data", "quality")}
stage = {name: True for name in ("accept", "reject", "fallback", "rollback-load")}
assert staged_release(ci, stage, "sha256:r9", "sha256:r9")["state"] == "ready"
assert staged_release(ci, {**stage, "rollback-load": False},
"sha256:r9", "sha256:r9")["state"] == "hold"
Performance and operating cost
Evaluating c CI and s staging results takes O(c + s) time and O(c + s) temporary set space. Full training and load tests dominate cost. Passing this simple gate does not verify report issuer identity or enforce atomic registry updates; those remain controls of the release system.
Common Mistakes
- Accepting quality evidence after the artifact bytes change.
- Skipping a rejected-input smoke request.
- Calling a canary safe before checking rollback compatibility.
- Treating a staging pass as proof of long-term production quality.
Read next
- ML tests: separate code, data, model and service failures
- Environment promotion: keep the tested artifact and contract together
- Model promotion: require evidence before changing the serving pointer
- Project: promote a receipt model with artifact and rollback evidence
- Project: release a versioned receipt feature admission gate
