A draft-critique-revise loop has value when the critique checks a fixed rubric or external facts that the first draft may have missed. Ask the critic for a defect category, location, evidence, and proposed check. A generic request to improve the answer often produces style changes without repairing a decision. Cap revisions, retain each candidate and critique, and rerun the original acceptance checks after every change. A model reviewing its own output is not independent ground truth; deterministic validators and human adjudication remain necessary for disputed cases.
Critique loops: require a named defect and a stopping rule
Decision in practice
An assistant drafts a maintenance handoff for pump PM-742. The critic finds a specific defect: the draft says pressure was measured by a sensor, while the record shows it was reported by a caller. The revision changes the attribution and keeps the pressure value linked to call segment CA-58. A second critique finds no remaining contract violation, so the loop stops. In a failure case, the critic only says 'make it more professional'; that does not authorize another expensive model call because it names no correctness defect.
Draft artifact: handoff H-742.
Critique fields: defect_type, affected_claim, evidence_id, required_check.
Valid defect: attribution mismatch at pressure claim; compare CA-58.
Revision budget: one pass; rerun attribution and evidence gates.
Stop: no named defect or budget exhausted -> accept or review by policy.Performance and operating cost
Each extra critique and revision call increases latency and token cost. With at most M revision rounds, model work is O(M) calls; validation over L output characters per round is O(ML). Test whether the loop repairs known defects without creating new ones. Record the proportion of critiques that name a real error and the proportion of revisions that pass after repair. If the critic and author share the same blind spot, repeated self-review can reinforce it. Use an independent case label or reviewer on high-risk decisions.
Common Mistakes
- Do not ask for endless refinement without a defect and a budget.
- Do not treat self-critique as an independent validator.
- Do not accept a revised answer without rerunning the original gates.
Connected lessons
- Prompt patterns
- Prompt Engineering
- Structured outputs: repair format without changing the decision
- Metamorphic tests: verify behavior when harmless details change
- Reasoning summaries: show checkable grounds, not invented certainty
- Multiple candidates: filter invalid answers before ranking
- Few-shot examples: teach the boundary with near misses
- Context distillation: shorten input without losing governing exceptions
- Negative controls: test the answer that should not be produced
- Project: verify a contract-renewal prompt at the boundary
- Reasoning and evidence checks
