Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Contrastive text cases: minimal edits and gold-label contracts

Last updated: 6 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A contrast case changes a small, reviewed part of a text to test a specific model decision. Store the edit, scope and expected outcome as separate evidence.

Specify the decision under test

A model may route “refund approved” correctly but also route “refund not approved” as approval. A contrast case changes one local condition so reviewers can inspect the decision boundary. Not every edit should flip the label: punctuation, harmless word order or a reviewed paraphrase may need the same outcome. Record task, base case ID, changed span, edit operator, original and edited gold labels, reviewer and rationale. Negation scope helps explain why a small token change can reverse an operational instruction.

Control what else changes

A negation test should not also change service, environment and timestamp. Keep protected identifiers and evidence spans fixed unless the test explicitly concerns them. A changed number might cross a policy threshold; merely replacing 47 with 82 does not guarantee a label flip unless the policy states that boundary. Review both texts independently to avoid anchoring the new label on the original. Quantity extraction preserves the value and unit needed for threshold cases.

Keep provenance and ambiguity

Store immutable base and edited text, offsets in the original representation and the annotation policy version. If the revised sentence is ambiguous, mark it unresolved rather than forcing a label. A reviewer can reject an edit if it accidentally changes two claims or creates impossible context. Cross-language or OCR edits need their own offset and normalization policies. Offset annotation keeps a reproducible link from the review record to the changed characters.

Measure pairs, not just examples

Ordinary accuracy can hide a classifier that gives the same answer to both sides of a contrast. Report base accuracy, edited accuracy and whether the relationship between predictions matches the relationship between gold labels. Slice by operator: negation, quantity, entity swap, temporal shift and paraphrase. The pairwise audit interprets those measures; the project applies them before release.

Implementation

python
def reviewed_edit(base_case, old_span, new_span, edited_label, reviewer_id):
    source_text = base_case["text"]
    if not old_span or source_text.count(old_span) != 1 or not reviewer_id:
        raise ValueError("edit span must be unique and reviewed")
    changed_text = source_text.replace(old_span, new_span, 1)
    if changed_text == source_text:
        raise ValueError("edit did not change the text")
    return {"base_id": base_case["case_id"], "base_label": base_case["label"],
            "edited_text": changed_text, "edited_label": edited_label,
            "old_span": old_span, "new_span": new_span,
            "reviewer_id": reviewer_id}

base = {"case_id": "runbook-47", "text": "Rollback is approved for gateway west.",
        "label": "approved"}
contrast = reviewed_edit(base, "is approved", "is not approved",
                         "blocked", "reviewer-82")
assert contrast["edited_text"] == "Rollback is not approved for gateway west."
assert contrast["base_label"] != contrast["edited_label"]

Performance and operating cost

Counting and replacing one span scan a text of n characters, so construction is O(n) time and O(n) space for the new string. Human review dominates cost when edits must preserve realistic context and verified labels. Unique substring matching is a small demonstration; production annotation should store explicit character spans and validate them against immutable text revisions.

Common Mistakes

  • Assuming every minimal edit should reverse the label.
  • Changing multiple facts while calling the case a single-variable test.
  • Inferring the new gold label from model output.
  • Losing the base text and annotation-policy version after an edit.

Read next

ai-data
natural-language-processing
Storage details