A sparse vectorizer and linear classifier form a measurable baseline when vocabulary and document-frequency statistics are fitted only on training text.
Sparse text baselines: fit vocabulary inside the training boundary
Use a baseline that can fail clearly
Represent support tickets with word or character features, then fit a linear classifier. Character n-grams can tolerate misspellings and invoice fragments; word features can be easier to inspect. Compare both only under the same split and metric. Do not jump to a larger model before knowing whether the task is separable with a modest representation. The split] determines whether a baseline score is credible.
Keep fitting inside the fold
Vocabulary selection and inverse-document-frequency weights learn from text. Fitting them on the whole corpus before splitting leaks information from evaluation messages. Put vectorization and classification in one pipeline, and fit only on training records within each fold. At serving time, transform with the saved vectorizer; do not rebuild a vocabulary from requests.
Interpret sparse evidence cautiously
A strong weight on a product code may reflect a useful routing rule or a temporary data artifact. Inspect high-weight terms and error cases by product and time slice. Short messages can become all-zero vectors if the token rule discards every term; define a fallback or abstain path. Slice evaluation] exposes that failure.
Record the comparison
Save vectorizer parameters, label order, split IDs and metric counts with the model. Compare against a majority-class and simple rules baseline. Report per-class precision and recall with denominators; a single accuracy number can hide a rare urgent queue.
Implementation
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import make_pipeline
ticket_router = make_pipeline(
TfidfVectorizer(analyzer="char_wb", ngram_range=(3, 5), min_df=2),
LogisticRegression(max_iter=800, class_weight="balanced"),
)
ticket_router.fit(training_texts, training_labels)
validation_labels = ticket_router.predict(validation_texts)Performance and operating cost
With N documents and Z nonzero features, sparse fitting scales mainly with Z and classifier iterations; dense conversion can require O(NV) memory for vocabulary size V. Measure matrix sparsity before changing models.
Common Mistakes
- Do not fit vocabulary or IDF weights before splitting.
- Do not rebuild the vectorizer at inference.
- Do not report accuracy alone for a rare urgent class.
