Apply the Natural Language Processing curriculum to a checked project.
Project brief
Required concepts
- Text corpus contracts: identity, label timing and annotation rules
- Unicode and tokenization: preserve meaning at the text boundary
- Entity spans: align annotations to the original text
- Sparse text baselines: fit vocabulary inside the training boundary
- Text validation: split conversations, duplicates and time together
- Text classification evaluation: inspect slices and allow abstention
- Text inference: package tokenizer, labels and reject paths
Completion standard
Submit runnable work, a data or run manifest, measured results and failure-case evidence.
