Self-consistency samples several candidate solutions to the same bounded problem, normalizes their final answers, and selects an answer that appears repeatedly. The method is useful when different valid routes can converge on one result. It does not make the candidates independent: the same model, prompt, and misleading input can produce the same wrong answer many times. Preserve the answer and the evidence needed to check it, rather than requiring a long private reasoning transcript. Where a deterministic check exists, run that check on the selected result before release.
Self-consistency: sample answers, then verify the winner
Operational case
A freight team must calculate the remaining load for shipment SH-731. The manifest says 47 sealed crates were booked, 19 were loaded at Dock C, and 11 were loaded at Dock F. Five sampled responses return 17, 17, 17, 28, and 17. The vote favors 17, yet the release gate still computes 47 minus 19 minus 11 and checks that neither dock entry was duplicated. If the manifest itself is stale, all five answers can agree on the wrong business fact. The system therefore attaches the manifest revision and refuses to present the result as current when that revision is missing.
Shipment SH-731: booked=47; dock_c=19; dock_f=11.
Candidates: 17 | 17 | 17 | 28 | 17.
Normalized winner: 17; vote share=4/5.
Independent arithmetic: 47 - 19 - 11 = 17.
Release only with a current manifest revision and no duplicate dock event.Performance and operating cost
For N candidates, generation cost and token use rise roughly with N; answer normalization and a frequency table take O(N) time and O(U) space for U distinct answers. Parallel calls can reduce wall-clock time but not total compute. Stop sampling once a predefined budget is reached, and compare the incremental accuracy with a single candidate plus verifier on a held-out set. Voting over free-form prose is unstable unless equivalent answers are normalized. Do not send more candidates merely to drown out a known data-quality problem.
Common Mistakes
- Do not call a four-of-five vote a correctness guarantee.
- Do not count differently worded versions of one answer as distinct outcomes.
- Do not skip the independent arithmetic or source-revision check.
Connected lessons
- Prompt patterns
- Prompt Engineering
- Multiple candidates: filter invalid answers before ranking
- Numeric prompts: let code calculate and the model explain
- Evaluation sets: measure the failure cases that matter
- Tree of Thoughts: branch only where a decision can be checked
- ReAct: alternate tool actions with checked observations
- Least-to-most prompting: solve smaller dependencies first
- Program-aided prompting: make computation executable and bounded
- Project: select and verify a reasoning pattern
- Reasoning patterns and operating limits
