Fairness evaluation begins with the decision rule and an appropriate set of comparable cases. Test whether irrelevant identity cues, dialect, or format changes alter the decision or explanation. Where group labels are sensitive, control access and use them only for a justified evaluation purpose. A disparity metric does not prove the prompt caused the gap; input quality, retrieval coverage, and policy may differ. Investigate the full path before changing wording, and have domain reviewers inspect consequential cases.
Fairness checks: test equivalent cases across groups and wording
Decision in practice
A rental-support assistant classifies maintenance requests by urgency. Two reports describe the same water leak and risk, but one uses formal English and one uses short colloquial phrases. Both should route to the same urgency tier and preserve the stated evidence. The team tests 58 paired reports with varied wording and recorded ground truth. When the informal reports are downgraded, it checks extraction and retrieval traces before changing the prompt. A reviewer verifies that any new examples do not encode the irrelevant wording pattern as a hidden priority rule.
Paired cases: same hazard, location, and timestamp; wording varied.
Expected relation: same urgency and evidence IDs.
Inspect: extraction fields, retrieved rule, decision, explanation.
Escalate: any high-consequence mismatch for domain review.Performance and operating cost
Paired testing at N base cases and V variants needs O(NV) model calls plus review. Small slices have noisy rates, so report counts and uncertainty rather than a grand fairness score. A test can detect sensitivity without identifying whether the prompt, retrieval, or source data caused it. Keep intervention records and re-run the paired cases after a fix. Do not remove a legitimate policy factor merely because it correlates with a group; evaluate relevance to the actual rule.
Common Mistakes
- Do not interpret an aggregate score as equal treatment.
- Do not infer cause from one disparity metric.
- Do not store sensitive group labels in general request logs.
Connected lessons
- Prompt Engineering
- Production prompt engineering
- Metamorphic tests: verify behavior when harmless details change
- Multilingual prompts: test policy meaning across languages
- Model migration: replay contracts before changing providers or versions
- Project: regression-test a customer triage prompt
- Advanced prompt engineering decisions
Apply this boundary: Decision thresholds: choose an action from probabilities and error costs.
