Evaluate a context-aware moderation queue that preserves quoted reports, routes uncertain cases and records corrected decisions.
Project: moderate a support thread with quotes and appeal review
Define the queue
A support thread contains an abusive original message, a reply quoting it to report the issue and a later apology. Moderation must identify the author and target of each act without punishing the reporter for the quoted span. Add a code-switched message with unclear intent and a case lacking the prior turn. Contextual labels keep the message, speaker and quote boundaries separate.
Build a reviewed fixture
Annotators view a bounded conversation window, mark target and quotation spans, and choose policy-specific labels or unresolved. Include disagreements and final adjudication reasons. Freeze the fixture by policy version before fitting thresholds. Keep raw messages restricted and strip unnecessary personal details from evaluation exports. A random message split can place near-identical quoted text across training and test, so group messages by thread and copied source.
Route and appeal
Score each message, then apply context-completeness and calibrated routing rules. A temporary hold must enter a timed reviewer queue; an appeal can reverse the decision while preserving the initial record. Do not treat a low model score as proof that a contextual threat is absent. Routing policy defines review and correction; the code below checks that held cases have a review record before closure.
Gate a policy release
Report missed targeted abuse, false holds on reports, review latency, appeal overturns and differences by language slice. Inspect individual errors rather than optimizing a single aggregate score. Release only after the quote-reporting and missing-context cases meet the reviewed policy. Keep a rollback route if new thresholds overload reviewers or wrongly hold a protected class of reports.
Implementation
def close_moderation_case(case):
if case["route"] in {"human-review", "temporary-hold-and-review"}:
if not case.get("reviewer_id") or not case.get("review_decision"):
return {"state": "open", "reason": "review-required"}
if case.get("appeal_state") == "pending":
return {"state": "open", "reason": "appeal-pending"}
return {"state": "closed", "effective_decision":
case.get("appeal_decision") or case.get("review_decision")
or case.get("initial_decision")}
case = {"route": "temporary-hold-and-review", "initial_decision": "hold"}
assert close_moderation_case(case)["reason"] == "review-required"
case.update(reviewer_id="moderator-47", review_decision="restore")
assert close_moderation_case(case) == {
"state": "closed", "effective_decision": "restore"}
Performance and operating cost
The final closure check is O(1) time and space for a fixed case record. Human review time, queue wait and appeal handling dominate operating cost. The example assumes the record has already been authenticated and its audit events persisted; production code must enforce those boundaries before changing visible content.
Common Mistakes
- Closing a held case without a reviewer decision.
- Treating a quoted complaint as the complaint author’s abuse.
- Dropping an appeal because an initial score was high.
- Reporting only aggregate accuracy while false holds cluster in one slice.
