Evaluate whether predictions change when meaning changes and remain stable when meaning does not. Pair-level results reveal failures hidden by aggregate accuracy.
Contrast-pair consistency: sensitivity, invariance and failure types
Separate two promises
A sensitivity case changes a gold decision: inserting “not” may turn approval into a block. An invariance case changes surface form while preserving the reviewed decision: “roll back” and “rollback” might be equivalent in a specific policy context. Both require human-checked labels. A model that flips every prediction appears sensitive but fails invariance; a model that never flips appears stable but misses changed meaning. Contrast construction records the intended relation for each pair.
Score the whole pair
Report correctness on each side and relation correctness for the pair. Relation correctness alone is insufficient: both predictions can be wrong yet preserve a same-label relation. Conversely, if the gold labels differ and only one prediction is right, the model may still flip. Maintain separate counts for base-only error, edited-only error, both wrong with expected relation and relation failure. This makes remediation concrete: collect better negation examples, fix threshold interpretation or revisit annotation policy.
Inspect correlated cases
Several edits from one base document are not independent observations. Aggregate by base case or incident as well as by edit, and keep train/test family boundaries. A large batch of trivial paraphrases should not drown out rare negation failures. Compare per-operator scores and confidence changes; a correct label with confidence collapse may still trigger a review or abstention path. Abstention policy should be tested on both members, not only on easy originals.
Use failures as regression fixtures
Pin approved contrast cases by corpus revision and model version. Re-run them after data, model, prompt or retrieval changes. A failure can arise from retrieval returning an old passage rather than the classifier itself; preserve retrieved evidence and intermediate outputs for diagnosis. The release project holds a deployment if a policy-critical contrast pair fails, even when average accuracy rises.
Implementation
def pair_report(contrast_pairs):
correct_sides = 0
correct_relations = 0
for pair in contrast_pairs:
correct_sides += pair["base_prediction"] == pair["base_label"]
correct_sides += pair["edited_prediction"] == pair["edited_label"]
gold_same = pair["base_label"] == pair["edited_label"]
predicted_same = pair["base_prediction"] == pair["edited_prediction"]
correct_relations += gold_same == predicted_same
return {"side_accuracy": correct_sides / (2 * len(contrast_pairs)),
"relation_accuracy": correct_relations / len(contrast_pairs)} if contrast_pairs else {
"side_accuracy": None, "relation_accuracy": None}
pairs = [{"base_label": "approved", "edited_label": "blocked",
"base_prediction": "approved", "edited_prediction": "approved"},
{"base_label": "approved", "edited_label": "approved",
"base_prediction": "approved", "edited_prediction": "approved"}]
assert pair_report(pairs) == {"side_accuracy": 0.75,
"relation_accuracy": 0.5}
assert pair_report([]) == {"side_accuracy": None, "relation_accuracy": None}
Performance and operating cost
Scoring p pairs takes O(p) time and O(1) additional space when only aggregate counters are retained. Diagnostic slices and per-base aggregation require O(p) stored outcomes. The example treats all pairs equally; a release gate should weight policy-critical operators explicitly and report their counts rather than hide them in a single mean.
Common Mistakes
- Calling a same-prediction pair correct when the gold decision changes.
- Using relation accuracy alone when both sides can be wrong.
- Counting many edits of one incident as independent test incidents.
- Ignoring a critical negation failure because overall accuracy improved.
Read next
- Contrastive text cases: minimal edits and gold-label contracts
- Project: gate a runbook answer release with contrast cases
- Text classification evaluation: inspect slices and allow abstention
- Textual entailment: premise, hypothesis and evidence boundaries
- Text validation: split conversations, duplicates and time together
