Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Contrast-pair consistency: sensitivity, invariance and failure types

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Evaluate whether predictions change when meaning changes and remain stable when meaning does not. Pair-level results reveal failures hidden by aggregate accuracy.

Separate two promises

A sensitivity case changes a gold decision: inserting “not” may turn approval into a block. An invariance case changes surface form while preserving the reviewed decision: “roll back” and “rollback” might be equivalent in a specific policy context. Both require human-checked labels. A model that flips every prediction appears sensitive but fails invariance; a model that never flips appears stable but misses changed meaning. Contrast construction records the intended relation for each pair.

Score the whole pair

Report correctness on each side and relation correctness for the pair. Relation correctness alone is insufficient: both predictions can be wrong yet preserve a same-label relation. Conversely, if the gold labels differ and only one prediction is right, the model may still flip. Maintain separate counts for base-only error, edited-only error, both wrong with expected relation and relation failure. This makes remediation concrete: collect better negation examples, fix threshold interpretation or revisit annotation policy.

Inspect correlated cases

Several edits from one base document are not independent observations. Aggregate by base case or incident as well as by edit, and keep train/test family boundaries. A large batch of trivial paraphrases should not drown out rare negation failures. Compare per-operator scores and confidence changes; a correct label with confidence collapse may still trigger a review or abstention path. Abstention policy should be tested on both members, not only on easy originals.

Use failures as regression fixtures

Pin approved contrast cases by corpus revision and model version. Re-run them after data, model, prompt or retrieval changes. A failure can arise from retrieval returning an old passage rather than the classifier itself; preserve retrieved evidence and intermediate outputs for diagnosis. The release project holds a deployment if a policy-critical contrast pair fails, even when average accuracy rises.

Implementation

python
def pair_report(contrast_pairs):
    correct_sides = 0
    correct_relations = 0
    for pair in contrast_pairs:
        correct_sides += pair["base_prediction"] == pair["base_label"]
        correct_sides += pair["edited_prediction"] == pair["edited_label"]
        gold_same = pair["base_label"] == pair["edited_label"]
        predicted_same = pair["base_prediction"] == pair["edited_prediction"]
        correct_relations += gold_same == predicted_same
    return {"side_accuracy": correct_sides / (2 * len(contrast_pairs)),
            "relation_accuracy": correct_relations / len(contrast_pairs)} if contrast_pairs else {
                "side_accuracy": None, "relation_accuracy": None}

pairs = [{"base_label": "approved", "edited_label": "blocked",
          "base_prediction": "approved", "edited_prediction": "approved"},
         {"base_label": "approved", "edited_label": "approved",
          "base_prediction": "approved", "edited_prediction": "approved"}]
assert pair_report(pairs) == {"side_accuracy": 0.75,
                             "relation_accuracy": 0.5}
assert pair_report([]) == {"side_accuracy": None, "relation_accuracy": None}

Performance and operating cost

Scoring p pairs takes O(p) time and O(1) additional space when only aggregate counters are retained. Diagnostic slices and per-base aggregation require O(p) stored outcomes. The example treats all pairs equally; a release gate should weight policy-critical operators explicitly and report their counts rather than hide them in a single mean.

Common Mistakes

  • Calling a same-prediction pair correct when the gold decision changes.
  • Using relation accuracy alone when both sides can be wrong.
  • Counting many edits of one incident as independent test incidents.
  • Ignoring a critical negation failure because overall accuracy improved.

Read next

ai-data
natural-language-processing
Storage details