The threshold for a rare specialist queue should reflect its error cost. A label-tree change is a data migration, not just a UI rename.
Hierarchical routing: branch thresholds and taxonomy migrations
Use branch-level decisions
A broad Billing score may be reliable while the Refund versus Chargeback distinction remains uncertain. Emit Billing and request review instead of inventing a leaf. Set thresholds using held-out reviewed tickets and the cost of wrong specialist routing. A global cutoff can over-route a rare leaf or reject too many common cases. Keep the model score, threshold bundle and final label reason with each decision.
Treat split and merge as migrations
When a queue splits, old examples may map to several new leaves; do not relabel them automatically from the former parent. When two labels merge, preserve the historical IDs and policy used at the time. Stage a migration with affected row counts and a reviewed sample. The current taxonomy must be acyclic and have known parents. Ancestor closure supplies the serving invariant.
Audit several kinds of error
Score branch precision, leaf precision, ancestor coverage, parent-only abstentions and actual routing outcomes. Compare latency and reviewer minutes. An apparent exact-path gain can come from predicting only common parents; a leaf gain can hide worse top-level routing. Slice by language and channel, and group repeat customer events across splits. Multilingual calibration affects branch thresholds.
Release with rollback
Version model, tokenizer, taxonomy graph, label display names and thresholds as one bundle. Replay a fixed audit when any part changes. Shadow-route traffic, inspect new branch conflicts and retain the old bundle for rollback. Update reports with a clear mapping from historical to current IDs rather than rewriting the past. The project exercises the staged migration.
Implementation
def route_hierarchy(parent_score, child_scores, parent_cutoff,
child_cutoffs):
if parent_score < parent_cutoff:
return {"state": "review", "labels": []}
accepted = [label for label, score in child_scores.items()
if score >= child_cutoffs[label]]
return {"state": "proposed", "labels": accepted or ["billing"]}
decision = route_hierarchy(0.91, {"refund": 0.62, "chargeback": 0.41},
0.75, {"refund": 0.82, "chargeback": 0.86})
assert decision["labels"] == ["billing"]
Performance and operating cost
Checking c child scores takes O(c) time and O(c) output space. The code illustrates one branch; production routing must apply the graph and thresholds to every eligible branch. Human review cost grows where specialist decisions are uncertain, so monitor parent-only rate and wrong-leaf rate together.
Common Mistakes
- Applying one threshold to every specialist leaf.
- Rewriting old labels after a taxonomy split without review.
- Measuring only exact path accuracy and hiding wrong top-level routes.
- Serving a new taxonomy with model outputs trained on old IDs.
