Tree of Thoughts treats a candidate intermediate plan as a node, expands several possible next steps, scores partial plans, and retains a bounded set for another round. Use it when an early choice changes later options and a single straight-line answer is brittle. A model-generated plausibility score is weak evidence; pair it with feasibility checks from the actual task. State the maximum branching factor, search depth, retained beam width, and stop rule before the search starts. Output a selected plan with its checked constraints, not a pile of speculative branches presented as facts.
Tree of Thoughts: branch only where a decision can be checked
Operational case
A warehouse needs a dispatch sequence for three temperature-controlled pallets and two refrigerated vans. The first candidate sends both heavy pallets in one van, violating its rated capacity. Another leaves a pallet beyond its cold-window limit. The search keeps only partial schedules whose cumulative weight and elapsed time remain feasible. At depth three it has two viable schedules, but one requires a driver who is unavailable; that branch is pruned from the verified roster. If the roster cannot be fetched, the system reports a conditional plan and requests dispatch approval instead of inventing availability.
Depth limit: 3 decisions; branch factor: at most 3; beam: 2 plans.
Check each node: van weight <= rated load; cold time <= 74 min.
Prune: over-capacity route; expired cold window; unavailable driver.
Stop: two verified feasible plans or 18 evaluated nodes.
Deliver: chosen schedule, constraints checked, unresolved roster status.Performance and operating cost
An unrestricted search with branch factor B and depth D can generate O(B^D) nodes. A beam of width W limits retained state to O(W), while expansions are about O(DWB), plus scorer and external-check costs. Bounded search still costs several model calls and can miss a good path after premature pruning. Record which constraint rejected each branch so a reviewer can distinguish a hard failure from a subjective score. Prefer direct deterministic scheduling when the whole problem can be solved by a trusted optimizer.
Common Mistakes
- Do not expand branches without a call and node budget.
- Do not rank a feasible route below an impossible route because its prose sounds less polished.
- Do not treat a model score as proof that a constraint was checked.
Connected lessons
- Prompt patterns
- Prompt Engineering
- Prompt decomposition: split stages at verifiable handoffs
- Multiple candidates: filter invalid answers before ranking
- Tool loops: set budgets, state checks, and a stopping condition
- Self-consistency: sample answers, then verify the winner
- ReAct: alternate tool actions with checked observations
- Least-to-most prompting: solve smaller dependencies first
- Program-aided prompting: make computation executable and bounded
- Project: select and verify a reasoning pattern
- Reasoning patterns and operating limits
