One reference answer does not define every valid response. Evaluate factual claims, abstention and the work a reader can safely perform.
Generation evaluation: claim support, disagreement and regressions
Write a task-specific rubric
Separate correctness, evidence support, completeness, unsafe instruction, clarity and appropriate abstention. A model may produce different valid wording from a reference; exact text similarity cannot establish whether an operational claim is supported. Review each claim against its cited passage and record unsupported, contradicted and unverified states. The claim contract makes the unit of evaluation explicit.
Build hard cases
Include stale runbook versions, near-identical error codes, missing documents, conflicting notes and permission-filtered passages. Hold entire incidents and document revisions out of prompt or model tuning. Evaluate short unanswered replies as legitimate outputs when evidence is absent. Sample production failures by risk and by ordinary traffic; only reviewing spectacular errors hides common low-grade omissions.
Calibrate reviewers and tools
Ask at least two reviewers to judge a subset without seeing each other’s decision. Adjudicate disagreement under a written policy. Automated checks can catch missing passage IDs or changed amounts, but they are screens rather than an oracle for support. Track judge version if a model assists review and compare its disagreements with human decisions. A rising acceptance rate may reflect reviewers becoming less strict rather than a better generator.
Compare releases on the same workload
Report unsupported claims per answer, dangerous action rate, useful-answer rate, unanswered rate and reviewer minutes. Slice by document age, language, question type and access scope. A release passes only if severe errors stay below the agreed ceiling and task utility improves. Keep an old approved bundle available for rollback. The project turns the rubric into a release test.
Implementation
def summarize_generation_audit(cases):
if not cases:
raise ValueError("audit needs reviewed cases")
unsupported = sum(case["unsupported_claims"] for case in cases)
dangerous = sum(bool(case["dangerous_action"]) for case in cases)
useful = sum(bool(case["useful"]) for case in cases)
return {"unsupported_per_answer": unsupported / len(cases),
"dangerous_action_rate": dangerous / len(cases),
"useful_answer_rate": useful / len(cases)}
audit = [{"unsupported_claims": 0, "dangerous_action": False, "useful": True},
{"unsupported_claims": 1, "dangerous_action": False, "useful": False}]
assert summarize_generation_audit(audit)["unsupported_per_answer"] == 0.5
Performance and operating cost
Summarizing n reviewed cases is O(n) time and O(1) extra space. The expensive part is claim-level human review and maintaining a representative frozen audit. Fewer, higher-quality cases with difficult retrieval and action boundaries can reveal more than a large pile of near-duplicate easy questions.
Common Mistakes
- Using one reference wording as the sole measure of correctness.
- Letting a judge model grade its own family without human calibration.
- Hiding severe action errors inside a high mean usefulness score.
- Counting an appropriate unanswered response as an automatic failure.
Read next
- Generated answers: claim evidence and unanswered state
- Project: release evidence-checked runbook answers
- Verify summary claims, corrections and human-review triggers
- QA answerability: calibrate abstention and evidence quality
- Adjudicate text labels before they become training truth
Continue the workflow: Contrast-pair consistency: sensitivity, invariance and failure types.
