Synthetic variants can widen training coverage, but they cannot replace independent native-language evidence.
Low-resource augmentation: provenance, label drift and slice audit
Record origin before generation
Every generated sentence needs a parent example, transformation type, source language, policy version and reviewer decision. Keep all descendants of one parent in one split. Otherwise a paraphrase can sit in training while its near-identical parent sits in the audit, creating a false gain. Treat translation and back-translation as candidate generators, not proof that the original intent survived. Paraphrase equivalence explains why changed negation or amounts matter.
Screen label preservation
Check protected entities, numbers, negation, modality and action target before a synthetic variant gets a label. A fluent sentence can reverse “refund not received” into “refund received.” Native-language reviewers should inspect a stratified sample, especially rare intents and mixed-script cases. Reject or relabel a variant when the intent changes. Keep synthetic and human-authored counts separate so a large generated pool does not masquerade as broad real coverage.
Run ablations against the same audit
Train the baseline, transfer model and augmented model under the same held-out groups and time window. Report performance on untouched native-language messages, not only transformed examples. Break results out by dialect, script, channel and rare intent; compare false automation with manual review volume. If augmentation improves the dominant label but harms a critical rare request, stop the release or narrow the rule.
Maintain the data boundary
A new label policy invalidates inherited synthetic labels until they are reviewed. Delete descendants when a parent is removed under retention policy. Track what percentage of training rows derive from each source and cap repeated variants from one customer event. Active learning can direct the next human-label budget; the applied project measures whether that spending helps.
Implementation
def assign_grouped_split(records, held_out_parent_ids):
result = {"train": [], "audit": []}
for record in records:
parent_id = record.get("parent_id", record["record_id"])
destination = "audit" if parent_id in held_out_parent_ids else "train"
result[destination].append(record["record_id"])
return result
rows = [{"record_id": "case-47"},
{"record_id": "variant-47-a", "parent_id": "case-47"},
{"record_id": "case-82"}]
split = assign_grouped_split(rows, {"case-47"})
assert split["audit"] == ["case-47", "variant-47-a"]
Performance and operating cost
The split pass is O(n) average time and O(n) output space for n records. Generation and native review dominate cost; multiple synthetic children from one parent add little independent evidence. Measure quality gain per reviewed parent and per serving language, including the added manual-review queue, before scaling generation.
Common Mistakes
- Putting synthetic children and their parent in different splits.
- Assuming fluent translated text preserved the original label.
- Reporting only aggregate gains while a rare intent regresses.
- Keeping generated descendants after deletion of the source example.
Read next
- Low-resource NLP: label budgets and transfer boundaries
- Project: adapt a support-intent router with scarce local labels
- Paraphrase equivalence: make hard negatives change the decision
- Active learning for text: uncertainty, diversity and coverage
- Text validation: split conversations, duplicates and time together
