Feedback prompts should quote or identify the learner's relevant reasoning step, compare it with a verified task rule, and ask for a revision the learner can perform. Separate a wrong conclusion from a sound conclusion reached by weak reasoning. A model may invent a misconception because one phrase resembles a common error; require an evidence span from the learner's answer and allow 'insufficient evidence.' For code or arithmetic exercises, deterministic tests should check outcomes before the tutor interprets the trace. Keep feedback about the task and method rather than personality. A reviewer should inspect feedback on ambiguous, dialectal, or incomplete responses so the tutor does not penalize language style as a technical error.
Tutor feedback: identify the specific reasoning step
Operational case
The learner writes, 'REQ-47 should reserve once because every request with the same SKU is the same operation.' The final count is right, but the rule is wrong: a second request ID can legitimately reserve the same SKU again. The tutor points to the phrase 'same SKU,' asks the learner to compare REQ-47 with a new ID REQ-58, and requests a revised rule keyed to request identity. Another learner writes only 'one'; the tutor cannot assume they know about durable receipts, so it asks what happens after a process restart. Neither response is graded as a fully explained answer from the number alone.
Learner span: 'every request with the same SKU is the same operation.'
Verified rule: duplicate request ID returns its prior receipt; new ID may act.
Feedback: compare REQ-47 retry with new REQ-58 for the same SKU.
Revision request: state the key and the persisted check.
If answer is only 'one': ask for mechanism; do not infer it.Performance and operating cost
For N responses and R rubric rules, naive comparison can require O(NR) checks; indexing feedback by task criterion cuts repeated review work. Model generation is only one cost. Independent tests, educator spot checks, and revision turns can dominate for code tasks. Keep the learner span, rule ID, model feedback, and revised answer together so a reviewer can see whether the feedback was justified. Do not compute a precise mastery score from a single short response or from stylistic features unrelated to the objective.
Common Mistakes
- Do not praise a correct number if the stated rule would fail on a new request ID.
- Do not invent a misconception without a supporting response span.
- Do not use confidence or writing style as a substitute for a task-specific check.
Connected lessons
- Prompt engineering applications
- Prompt Engineering
- Evidence IDs: make generated claims auditable against supplied records
- Generated tests: verify the oracle before trusting coverage
- Classification prompts: write the label boundary first
- Tutor prompts: start from an observable learning objective
- Tutor hints: escalate help without leaking the answer key
- Assessment prompts: separate feedback from final grading
- Tutor release checks: minimize learner state and test transfer
- Project: review a retry-safety training tutor
- Tutoring and assessment prompt decisions
