Test simplified technical text for reader understanding, factual fidelity and action safety instead of relying on readability scores alone.
Text simplification evaluation: comprehension and release slices
Two objectives need two measurements
Ease of reading and preservation of meaning can pull in different directions. A notice can be easy to read because it omits the hard condition that mattered most. Ask readers to answer task-specific questions from both source and rewrite. Score factual answers and decision errors separately from speed or subjective ease. Keep an explicit “not stated” answer for missing information. Protected facts define what cannot disappear.
Build audience slices
Test readers who know the internal service names and readers who do not. Include screen-reader use, mobile viewing and second-language readers where the product serves them. Do not assume a single target reading level captures these needs. The same output may be clear to engineers and confusing to customers because an acronym remains unexplained. Review examples rather than treating an aggregate grade as a release certificate. Abbreviation scope identifies domain terms that require expansion.
Check edit-level harm
Annotate each rewrite operation: deletion, split, reorder, substitution and explanation. Flag edits that change actor, time, unit, certainty, negation or condition. Pair human checks with automatic invariants for identifiers and numbers. An automatic metric can surface candidates for review but should not silently approve notices that could alter an operational choice. NLI evidence contracts offer one way to compare source and generated claims.
Use a conservative release state
A candidate should move to publication review only when protected facts are present, comprehension questions retain their answers and an editor has checked the critical claims. Keep the source revision and reviewer decision. If a fact changes upstream, invalidate the simplified copy rather than assuming it remains true. The status notice project demonstrates that invalidation path.
Implementation
def compare_reader_answers(source_answers, rewrite_answers):
if set(source_answers) != set(rewrite_answers):
return {"state": "review", "reason": "question-set-mismatch"}
changed = [question for question in source_answers
if source_answers[question] != rewrite_answers[question]]
if changed:
return {"state": "review", "changed_questions": changed}
return {"state": "answers-preserved"}
source = {"which_service": "gateway-west",
"certainty": "possible", "when": "after 47 minutes"}
rewrite = dict(source)
assert compare_reader_answers(source, rewrite)["state"] == "answers-preserved"
assert compare_reader_answers(source, {**rewrite, "certainty": "certain"}) == {
"state": "review", "changed_questions": ["certainty"]}
Performance and operating cost
Comparing q question keys and answers takes expected O(q) time and O(q) space for key sets and changed questions. Human comprehension testing has recruitment and review cost. Matching answer labels is only a check on the selected questions; unasked meaning differences can remain, so critical notices still require editorial review.
Common Mistakes
- Reporting one readability number as meaning fidelity.
- Testing only experts when customers are the target readers.
- Omitting “not stated” from comprehension questions.
- Keeping a simplified copy after its source fact changes.
Read next
- Plain-language rewrites: protect facts before shortening text
- Project: rewrite an incident status notice without changing its claim
- Textual entailment: premise, hypothesis and evidence boundaries
- Abbreviations: definition scope, collisions and unknown forms
- Revision impact: invalidate derived claims, indexes and answers
