A text evaluation partition must separate repeated conversations and future time, or near-duplicate wording can inflate the result.
Text validation: split conversations, duplicates and time together
Identify repeat exposure
One customer may reopen a ticket by copying an earlier description. Templates may also create many near-identical messages. Random document splitting can place the same phrasing on both sides and make memorization look like generalization. Group by conversation or customer according to the deployment question, and audit high-similarity cross-partition pairs. Grouped validation] provides the broader model-selection rule.
Respect the prediction clock
A future validation window asks whether the classifier handles new products and language. Keep training before the cutoff, then apply an entity rule to prevent a customer conversation spanning both sides from leaking. When these constraints conflict, document excluded records and the resulting population. An evaluation with only historical random folds answers a different question.
Handle delayed labels
Some tickets receive a final category days after submission. Select evaluation examples by submission time, but calculate metrics only when labels have matured under a fixed cutoff. Otherwise recent difficult cases can disappear and the apparent quality rises. Delayed-label monitoring] uses the same maturity rule after launch.
Audit the split itself
Report counts by label, language, customer and month. Confirm disjoint group IDs and inspect a bounded near-duplicate sample. If the urgent category is absent from one fold, do not conceal it by averaging the fold metric. Revisit the split design or collect more examples.
Implementation
def assert_ticket_split(training_rows, evaluation_rows):
training_conversations = {row.conversation_id for row in training_rows}
evaluation_conversations = {row.conversation_id for row in evaluation_rows}
if training_conversations & evaluation_conversations:
raise ValueError("conversation crosses evaluation boundary")
latest_training = max(row.submitted_at for row in training_rows)
earliest_evaluation = min(row.submitted_at for row in evaluation_rows)
if latest_training >= earliest_evaluation:
raise ValueError("future overlap in text split")Performance and operating cost
Group and time checks are O(N) over ticket records with O(G) memory for G groups. Near-duplicate search is costlier than exact matching; blocking or approximate indexing avoids all-pairs O(N²) comparison.
Common Mistakes
- Do not scatter replies from one conversation across partitions.
- Do not score labels before their maturity cutoff.
- Do not hide missing classes in a fold average.
