A model can confuse a writing variety with a label. Audit the task decision by language variety without guessing a person’s identity.
Dialect-aware NLP audits: slices and label disagreement
Define the task before the slice
For support routing, the target is the customer’s request, not whether wording resembles a preferred register. Specify the intended label policy with examples in the writing varieties the product actually receives. Do not infer a customer’s race or identity from dialect markers. If a slice is needed for quality review, use consented or appropriately governed labels, and report sample sizes and uncertainty. Reviewer disagreement is part of the evidence.
Separate distribution from error
A different mix of intents can change aggregate accuracy even when the model makes the same conditional errors. Measure false positive and false negative rates for comparable labels within each sufficiently supported slice. Review minimal pairs where the operational intent stays fixed while wording varies, but do not claim a synthetic rewrite fully represents a community’s speech. Preserve original messages and source context for native-language review.
Examine the annotation process
Reviewers may interpret slang, reclaimed terms or terse code-switched messages differently. Record rationale, policy version and unresolved cases instead of collapsing every disagreement to a majority vote. A model trained on a disputed label inherits that policy decision. Ask qualified reviewers to examine borderline cases, and separate model errors from annotation ambiguity. Low-resource label budgets affect which slices can be assessed.
Choose a release response
Report per-slice precision, recall, reject rate, review burden and the largest error gaps with intervals. If one slice lacks evidence, route uncertain cases to review rather than declare parity. A threshold change can reduce false positives while increasing missed requests; make that tradeoff explicit. The release gate controls the data boundary, and the project tests it.
Implementation
def slice_error_counts(records, slice_name):
selected = [row for row in records if row["audit_slice"] == slice_name]
if not selected:
return {"supported": False, "count": 0}
false_positive = sum(row["predicted"] and not row["reviewed"]
for row in selected)
false_negative = sum(not row["predicted"] and row["reviewed"]
for row in selected)
return {"supported": True, "count": len(selected),
"false_positive": false_positive, "false_negative": false_negative}
audit = [{"audit_slice": "variety-k", "predicted": True, "reviewed": False},
{"audit_slice": "variety-k", "predicted": False, "reviewed": True}]
assert slice_error_counts(audit, "variety-k")["false_positive"] == 1
Performance and operating cost
Scanning n reviewed rows is O(n) time and O(n) temporary space for a selected slice in this direct implementation. Streaming counters would use O(1) extra space. The expensive part is obtaining trustworthy reviewed examples and resolving policy disagreements, not computing rates. Never publish a tiny slice as a stable performance estimate.
Common Mistakes
- Using dialect markers as a proxy for personal identity.
- Comparing pooled accuracy when intent mix differs across slices.
- Treating a synthetic spelling rewrite as complete dialect evidence.
- Hiding reviewer disagreement behind a single label.
