Review a candidate digest against intended use, slice evidence and production routing before any traffic is promoted.
Project: approve a receipt model with a scoped evidence dossier
Assemble one immutable dossier
Use the candidate receipt model digest, training run manifest, feature contract, evaluation snapshot and proposed response policy. State that the score prioritizes human review; it must not become an automatic payment denial. Name supported countries and currencies, the decision threshold owner and the safe route for unsupported inputs. The operating card supplies the human decision, while the artifact digest pins the executable bytes.
Probe three uncomfortable cohorts
Evaluate established merchants, new merchants and low-quality scans on the same mature outcome window. The candidate improves overall recall, but the new merchant cohort has only 28 mature outcomes against a 47-outcome minimum. Mark that slice insufficient. Do not turn the missing evidence into a favorable score. For low-quality scans, inspect the error count and feature availability together; a bad result may call for admission control rather than another threshold. Record eligible, scored, reviewed and labeled counts for each cohort.
Choose an enforceable release state
The sample dossier should hold broad promotion. It may propose a restricted pilot that routes new merchants to manual review, but only after an owner approves that compensating control and verifies the API enforces it. Store the control expiry and the evidence that will end the restriction. Slice gates produce pass, fail or insufficient; promotion attestation must not treat “insufficient” as a pass. Re-evaluate whenever scope or threshold changes.
Deliver reviewable artifacts
Publish the card revision, model digest, cohort query ID, metrics, maturity rule, gate results, limitations, compensating control and reviewer decision. Include an operator test that sends an unsupported currency through the endpoint and observes manual review. A dashboard link alone is weak evidence: it can change after label corrections or data retention. Preserve a dated decision record and tell registered consumers which fields and routes they may expect.
Implementation
def approve_receipt_release(slices, scope_route_verified):
if not scope_route_verified:
return {"state": "hold", "reason": "scope-route-unverified"}
failed = [name for name, state in slices.items() if state == "fail"]
unknown = [name for name, state in slices.items()
if state == "insufficient"]
if failed:
return {"state": "hold", "reason": "slice-failure", "slices": failed}
if unknown:
return {"state": "restricted-review", "slices": unknown}
return {"state": "candidate"}
review = {"established": "pass", "new-merchant": "insufficient",
"low-quality-scan": "pass"}
assert approve_receipt_release(review, True)["state"] == "restricted-review"
assert approve_receipt_release(review, False)["state"] == "hold"
Performance and operating cost
The decision scan is O(s) time and O(s) output space for s slices. Gathering evidence dominates: mature outcomes can take weeks and route verification needs integration testing. A restricted decision is not production approval by itself; the proposed compensating route must actually work, have an owner and expire or be re-reviewed.
Common Mistakes
- Promoting because the total metric improved while a material slice is unknown.
- Stating manual review as a limitation without exercising the route.
- Saving metrics without the evaluation snapshot or model digest.
- Leaving a restricted approval open indefinitely.
Read next
- Model cards as operating contracts: scope, evidence and limits
- Slice quality gates when labels are sparse or delayed
- Promotion evidence: bind evaluation, contract and rollback to one digest
- Project: release a versioned receipt feature admission gate
- Project: promote a receipt model with artifact and rollback evidence
Continue the workflow: Project: audit an explanation for a disputed receipt decision.
