Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Project: build a checked synthetic return-triage suite

Last updated: 2 Oct 202619 min read
project
AdvancedBy AITrove Editorial

Build a release evaluation kit for a fictional return-triage assistant. The assistant may answer a routine delivered-scan case, send a missing-scan dispute for review, or report that evidence is unavailable. Its tool layer separately verifies account scope and effect permissions. Author cases from this abstract rule contract, not from a customer ticket export. The target is 24 accepted fixtures across evidence states, ambiguity, hostile notes, and effect boundaries. Give each fixture a fictional ID, lineage key, mutation reason, policy version, expected action, and reviewer decision.

Generate and reject

Start with RT-641, a delivered-scan routine case. Mutate only the evidence field for RT-642 and RT-643; the missing scan routes to review, and unavailable evidence routes to unknown. Add a customer note that asks the assistant to approve a claim despite the scan. That note tests instruction quarantine and must not become policy. Reject a generated case that labels a missing scan routine. Use the Python oracle from the connected lesson to catch label contradictions and duplicate IDs; add a wider rule table before testing account permission. Build a slice matrix and reject color-only paraphrases that do not change the decision path.

Protect and hold out

Screen all 24 cases for copied identifiers and distinctive production wording. Quarantine any case whose fictional name still carries a real parcel reference or rare incident detail. Assign related mutations to the same development or release split. Lock a critical set containing a stale scan, a conflicting carrier event, and a hostile note mixed with a valid question. If that set is exposed during prompt tuning, record the exposure and replace the compromised holdout before claiming an independent score. Produce a report with accepted and rejected counts by reason, slice coverage, oracle version, privacy review, and prompt-release decision.

Output
RT-641: delivered_scan -> routine.
RT-642: missing_scan -> needs_review.
RT-643: unavailable -> unknown.
Reject: missing_scan -> routine; invalid oracle label.
Holdout: group seed siblings; record every exposure.

Performance and operating cost

For S seeds and M mutation operators, generation proposes O(SM) candidates. Validating C cases against a constant rule table is O(C); naive near-duplicate comparison can be O(C squared). Review cost dominates when a case is ambiguous, privacy-sensitive, or dependent on an old policy. Keep distinct counts for generated, structurally valid, label-valid, privacy-cleared, and holdout-eligible cases. Do not convert the acceptance fraction into a claim about live traffic or privacy guarantees. The release gate blocks a critical wrong action even when the average synthetic score improves.

Common Mistakes

  • Do not treat a generated label as ground truth without a separate oracle.
  • Do not confuse a fictional ID with a privacy guarantee.
  • Do not put close seed siblings on both sides of a held-out split.

Connected lessons

prompt engineering
synthetic evaluations
Storage details