A model card should tell a release reviewer where a model may run, what was measured and which uses remain unsupported.
Model cards as operating contracts: scope, evidence and limits
Name the decision, not just the algorithm
A receipt triage model assigns a review priority to eligible digital receipts. The card must state the decision owner, affected workflow, intended user, permitted geography and fallback when evidence is missing. A statement such as “detects fraud” is too broad: it hides who acts on the score and may invite use as an automatic denial. Keep the threshold and its owner alongside the decision description. Run lineage identifies the trained bytes; the card explains their supported use.
Connect every claim to a frozen cohort
Record the data snapshot, labeling rule, time window, exclusion criteria, threshold and mature-outcome coverage behind each metric. Report overall quality alongside named slices that matter to operations, such as new merchants, uncommon currencies or low-quality scans. A slice with few labeled outcomes is unknown, not proof of parity. Link the exact evaluation run and model digest through promotion evidence. Avoid a screenshot of a dashboard that cannot be reproduced after labels are corrected.
Write limitations as actions
A useful limitation names a trigger and response. If a receipt currency was absent from evaluation, route that case to manual review until evidence exists. If a feature goes stale, follow the freshness policy rather than silently score an old value. Separate unsupported use from known error patterns and unresolved risks. The card belongs to a versioned release: changing the model, threshold, eligible population or response policy may require a new review, even if the artifact digest is unchanged.
Keep the record alive after release
Assign a review date and owner. Add field failures, overridden decisions and corrected labels as dated entries without rewriting the original approval. Compare deployed population with the approved scope; sustained movement into an untested slice should open a fresh evaluation, not quietly broaden the card. Consumer inventory shows where the model is used, while the review project turns a vague claim into an enforceable release boundary.
Implementation
def check_model_scope(card, request):
if card["digest"] != request["model_digest"]:
return "hold:digest"
if request["country"] not in card["approved_countries"]:
return "manual-review:geography"
if request["currency"] not in card["evaluated_currencies"]:
return "manual-review:currency"
return "score"
card = {"digest": "receipt-r47", "approved_countries": {"IN", "SG"},
"evaluated_currencies": {"INR", "SGD"}}
receipt = {"model_digest": "receipt-r47", "country": "IN", "currency": "INR"}
assert check_model_scope(card, receipt) == "score"
assert check_model_scope(card, {**receipt, "currency": "EUR"}) == "manual-review:currency"
Performance and operating cost
Membership checks are O(1) expected time and O(c) memory for c allowed scope values. Maintaining the card costs review time and evidence storage, but the operational cost of a missing scope boundary is unmeasured use. The code checks a small gate; it does not measure whether the listed slices were evaluated well.
Common Mistakes
- Treating a single aggregate score as evidence for every population.
- Writing a limitation without a routing or escalation action.
- Approving a threshold change under the old card without review.
- Overwriting the original approval after a production incident.
Read next
- Slice quality gates when labels are sparse or delayed
- Project: approve a receipt model with a scoped evidence dossier
- Training manifests: link data, code, configuration and artifact
- Model promotion: require evidence before changing the serving pointer
- Model retirement: find consumers before removing a version
