A metamorphic test creates related inputs with a known relationship between their outputs. Harmless changes, such as swapping the order of two equivalent documents or renaming an irrelevant case token, should not change the governing decision. Material changes, such as removing the only eligibility clause, should change the result to unknown. This tests invariants even when a full answer key is difficult to author. Keep generated variants outside the prompt examples and inspect failures before treating them as new training cases.
Metamorphic tests: verify behavior when harmless details change
Decision in practice
A service-credit prompt approves a case when a signed agreement contains clause CL-27. The test suite copies the case and changes the customer token; approval should remain. It then places a harmless appendix before the clause; approval should remain with the same evidence ID. A third variant removes CL-27; approval must become review. If the output changes merely because the appendix moved, the team has found a position sensitivity defect rather than a new policy rule.
Base: signed clause CL-27 present -> approve.
Variant A: rename unrelated case token -> same decision.
Variant B: move appendix -> same decision and evidence ID.
Variant C: remove CL-27 -> review, no approval effect.Performance and operating cost
If N base cases each produce V variants, the suite makes O(NV) model calls before repeats. Favor transformations tied to known failure modes rather than random paraphrases with unclear expected behavior. Automated relation checks can compare enum decisions and evidence IDs, while humans review disputed semantic equivalence. Report both false invariance and false sensitivity: a system that never changes its answer can pass a weak same-output test while ignoring material evidence.
Common Mistakes
- Do not assume every paraphrase preserves meaning.
- Do not test only same-output relations.
- Do not add failed variants to few-shot examples before reserving a fresh holdout.
Connected lessons
- Prompt Engineering
- Production prompt engineering
- Evaluation sets: measure the failure cases that matter
- Long context: make inclusion and truncation testable
- Prompt caching: reuse stable context with a versioned expiry
- Human handoff: preserve evidence and the reason for uncertainty
- Project: regression-test a customer triage prompt
- Advanced prompt engineering decisions
Continue with: Critique loops: require a named defect and a stopping rule.
