Review a parcel-condition model with unknown annotation states, separate label cutoffs, case-level review workload, per-label errors and a sealed target period.
Multi-label parcel review project
Freeze the inspection and label contract
Define the parcel or incident unit, the allowed condition flags and which inspection establishes each negative. Keep unknown annotations distinct from confirmed absence. Group related photos and preserve a later target period. The schema lesson is the record of these decisions.
Build a transparent first model
Use a binary head per label on eligible training rows, then compare any joint-label alternative on the same grouped development cohort. Record feature timing and annotation coverage. The baseline shows what a joint model must improve.
Select the action policy
Choose per-label cutoffs using declared costs and an isolated development set. Combine triggered flags into unique parcel actions and check the actual review desk’s deadline and capacity. Do not choose cutoffs from final-test outcomes. Cost selection and review capacity are both needed.
Evaluate the right denominators
Report exact-set accuracy, Hamming loss, micro and macro F1, per-label support, false negatives, calibration and average true and predicted label count. These describe different properties; no one score is a release verdict. Check depot and camera slices and disclose selectively inspected labels. The metric guide and calibration guide define the reports.
Name the bounded next step
The code checks evidence fields in a release packet. It cannot decide whether measured misses or costs are acceptable. Attach the actual values, the owner, a rollback route and the next evaluation window; hold the pilot review if label coverage or capacity is still unknown.
Implementation
def multilabel_review_gate(packet):
required = {
"unknowns_preserved": "unknown labels",
"parcel_groups_sealed": "parcel groups",
"cutoffs_selected_on_development": "development cutoffs",
"per_label_support_reported": "label support",
"case_review_capacity_checked": "review capacity",
"calibration_audited": "label calibration",
"rollback_owner_named": "rollback owner",
}
missing = [label for field, label in required.items() if not packet.get(field)]
return "eligible for pilot review" if not missing else "hold: " + ", ".join(missing)
parcel_packet = {
"unknowns_preserved": True, "parcel_groups_sealed": True,
"cutoffs_selected_on_development": True,
"per_label_support_reported": False,
"case_review_capacity_checked": False,
"calibration_audited": True, "rollback_owner_named": True,
}
assert multilabel_review_gate(parcel_packet) == (
"hold: label support, review capacity"
)Performance and operating cost
The gate costs O(K) for K checks. Multi-label evaluation over N cases and L conditions is O(NL), while a T-threshold grid is O(LTN) in a direct implementation. The larger operating costs are complete annotation, review capacity and validation across parcel groups and time.
Common Mistakes
- Do not replace unknown target states with negatives.
- Do not release from a single pooled metric when a rare costly label fails.
- Do not sum per-label alert counts and mistake them for unique parcels requiring review.
Read next
- Multi-label schema and unknown targets
- Binary relevance and label dependence
- Per-label thresholds and action cost
- Multi-label metrics and denominators
- Multi-label calibration and label cardinality
- Label contracts: decision unit, taxonomy and review rules
Continue the workflow: Sequential stock-control policy review project.
