Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Text feedback loops: correction provenance and shadow audits

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Agent corrections can improve a model, but a production workflow also changes what gets reviewed and therefore what enters training.

Record the decision path

For each ticket, keep model proposal, confidence, policy version, agent action, correction reason and review time as separate events. An agent might leave an incorrect route untouched because the next queue handles it; “not corrected” is not the same as “correct.” If the model auto-routes easy cases, the reviewed pool becomes harder and no longer represents incoming traffic. Acquisition coverage helps design a review sample.

Reserve independent review

Randomly sample a small governed share of confident and abstained cases for independent labeling, in addition to normal agent corrections. Keep the fixed final audit out of training and threshold tuning. Group repeated customer events before splits. A shadow model should score the same intake records as production without changing their routes; compare it only after mature independent labels arrive.

Admit corrections carefully

A correction must carry source text revision and current label policy. Exclude labels created solely by a model’s previous decision or by a downstream queue that encodes the answer. Review taxonomy migrations and disputed cases before training. Weak-label audits provide a parallel provenance check. Delete training derivatives if the underlying ticket is removed under retention rules.

Compare operating outcomes

Track independently reviewed precision and recall, disagreement, abstention, queue age and rework. Evaluate old and shadow models on the same cohort. If a new model lowers manual reviews by routing risky cases incorrectly, the apparent efficiency gain is false. The drift project ties this audit to release and rollback decisions.

Implementation

python
def admit_agent_correction(record, current_policy):
    if record["policy_version"] != current_policy:
        return False
    if record["source_revision"] != record["reviewed_revision"]:
        return False
    if record["review_origin"] not in {"independent", "agent-correction"}:
        return False
    return record["reviewed_label"] is not None

correction = {"policy_version": "labels-r8", "source_revision": "note-r47",
              "reviewed_revision": "note-r47", "review_origin": "independent",
              "reviewed_label": "refund"}
assert admit_agent_correction(correction, "labels-r8")

Performance and operating cost

Admission is O(1) per reviewed record; maintaining an independent sample costs reviewer time. Shadow inference can temporarily double model calls, but it separates measured improvement from intervention effects. A small representative review stream is often worth more than a large pile of unverified production outcomes.

Common Mistakes

  • Treating “agent did not correct it” as a positive label.
  • Evaluating only cases the model chose to abstain from.
  • Training on model-generated labels without independent review.
  • Mixing policy versions after a taxonomy update.

Read next

Continue the workflow: Search feedback: position bias, query drift and audit design.

ai-data
natural-language-processing
Storage details