Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Synthetic case oracles: reject invalid labels before scoring a model

Last updated: 2 Oct 202611 min read
tutorial
AdvancedBy AITrove Editorial

An evaluation oracle states the expected action independently of the system being tested. For synthetic cases, write the invariant first: a delivered scan permits the routine path only when the policy and account checks also pass; a missing scan requires review; unavailable evidence requires an unknown answer. The small sample below isolates evidence-state routing and deliberately omits the wider account policy. Generate candidate cases, then run schema, identifier-uniqueness, and rule-consistency checks before any model score is calculated. A reviewer should approve ambiguous or policy-dependent labels. If the generator and judge share the same unsupported assumption, agreement between them is not evidence of correctness.

Operational case

A batch contains four fictional return tickets. RT-641 has a delivered scan and a routine label, RT-642 and RT-644 have missing scans and review labels, and RT-643 has unavailable evidence with an unknown label. An accidental fifth draft says 'missing_scan' but labels it routine. The oracle rejects it before the assistant is evaluated. A separate case can test whether an authorized account owner is present, but it needs a broader rule table and cannot be folded into this simple evidence-only function. The reviewer records which policy version defines each expected action so a later policy change does not silently rewrite past results.

python
from collections import Counter

synthetic_cases = [
    {"id": "RT-641", "evidence": "delivered_scan", "label": "routine"},
    {"id": "RT-642", "evidence": "missing_scan", "label": "needs_review"},
    {"id": "RT-643", "evidence": "unavailable", "label": "unknown"},
    {"id": "RT-644", "evidence": "missing_scan", "label": "needs_review"},
]
expected_label = {
    "delivered_scan": "routine",
    "missing_scan": "needs_review",
    "unavailable": "unknown",
}
case_ids = [case["id"] for case in synthetic_cases]
assert len(case_ids) == len(set(case_ids))
for case in synthetic_cases:
    assert case["label"] == expected_label[case["evidence"]]
counts = Counter(case["label"] for case in synthetic_cases)
print(f"oracle: {len(synthetic_cases)}/{len(synthetic_cases)} valid")
print(f"slices: routine={counts['routine']} review={counts['needs_review']} unknown={counts['unknown']}")

Performance and operating cost

For C cases and a constant-size evidence rule table, the sample validates labels in O(C) time and O(C) space for unique identifiers and label counts. A richer oracle may need joins against policy and account snapshots; its cost follows those lookups. The sample catches contradictory labels, not subtly unrealistic language or missing context. Keep the oracle's code and policy version separate from the prompt under test. When a reviewer changes an expected label, record the reason and rerun the score from the original case rather than silently editing a previous report.

Common Mistakes

  • Do not score a model against an unvalidated generated label.
  • Do not use the model under test as the sole oracle for its own cases.
  • Do not claim an evidence-only sample validates account permission or full return policy.

Connected lessons

prompt engineering
synthetic evaluations
Storage details