Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Multiple candidates: filter invalid answers before ranking

Last updated: 5 Oct 202610 min read
tutorial
AdvancedBy AITrove Editorial

Candidate sampling produces more than one proposed answer for the same frozen input. A majority vote can favor a shared mistake, especially when every candidate sees the same missing or misleading record. First reject candidates that fail schema, authorization, evidence, or task-specific checks; only then rank the survivors by a separate rubric. Keep the input, retrieved bundle, and model settings recorded. If no candidate passes, return review. Candidate diversity is an experiment, not a license to convert agreement into truth.

Decision in practice

Three candidates assess renewal RN-284. Two approve it using expired policy POL-63; one requests review using the current policy POL-64 but lacks a hold record. A vote would approve. The application instead rejects the two stale citations and preserves the review candidate. The reviewer then checks whether the missing hold can be retrieved. The selection test includes a run in which all candidates use expired records, so the correct result is no accepted candidate. The team compares this method with a single-candidate baseline before paying for extra generations.

python
current_evidence_ids = {"POL-64"}
candidates = [
    {"id": "RN-284-A", "decision": "approve", "evidence_ids": ["POL-63"]},
    {"id": "RN-284-B", "decision": "approve", "evidence_ids": ["POL-63"]},
    {"id": "RN-284-C", "decision": "review", "evidence_ids": ["POL-64"]},
]

accepted = [
    candidate for candidate in candidates
    if candidate["evidence_ids"]
    and set(candidate["evidence_ids"]) <= current_evidence_ids
]
print("accepted:", ", ".join(item["id"] for item in accepted) or "none")

Expected output: accepted: RN-284-C

Performance and operating cost

For K candidates and C cited IDs per candidate, this structural filter costs O(KC) expected time and O(K) output space, excluding the input. Generating K responses can cost roughly K times a single call before any ranking. The filter only verifies reference membership; it does not prove that POL-64 supports the stated decision. Add a claim-level review for high-impact outcomes. Track valid-candidate rate, accepted error rate, and total cost together. A method that merely chooses the most common wrong answer has increased expense without improving safety.

Common Mistakes

  • Do not use majority agreement as a substitute for evidence checking.
  • Do not rank a candidate that fails a hard gate.
  • Do not force selection when every candidate is invalid.

Connected lessons

Related implementation

Continue with: Self-consistency: sample answers, then verify the winner.

prompt engineering
evaluation
Storage details