Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Lemmas, part of speech and domain terms in a search index

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Word forms help recall, but careless reduction can erase a product name or make two different actions look identical.

Define the index identity

Store original text and offsets before making any searchable representation. A lemma is a normalized word form conditioned on a grammatical reading; a stem is a mechanical reduction and may not be a word. Neither should replace the original. In “the charge backs out” and “a chargeback arrived,” a nearby sequence of letters does not establish the same event. Keep exact identifiers, product names and quoted error text in separate fields. Unicode token boundaries are part of this contract.

Use context for ambiguous forms

A token such as “returns” may be a noun in a policy title or a verb in a customer report. Part-of-speech assignment influences a lemma and can fail on terse headings. Store the tagger version, language and confidence beside the derived form. Where uncertainty is high, index both the surface form and a small reviewed expansion set. Never rewrite the source sentence in place. A one-size suffix rule is especially brittle for mixed-language or technical vocabulary.

Protect terms with operational meaning

Maintain an allowlist for domain tokens such as SKU names, incident codes and command flags. A product name that resembles an ordinary plural is still an exact entity. Review each normalization change against exact-match queries and misspellings. Use script profiling before applying language-specific morphology; a message can switch language in the middle of a sentence. Version the term list and rebuild derived postings when it changes.

Evaluate retrieval effects

Compare exact search, surface-plus-lemma search and a conservative domain dictionary on held-out queries. Report recall at a fixed review depth, rank changes, false matches and exact-code preservation. A gain on broad “refunds” queries cannot excuse a lost search for order ZX-47. Hybrid ranking combines these lexical signals with semantic candidates, while the project tests a release in context.

Implementation

python
def searchable_forms(surface, part_of_speech, lemma_table, protected):
    exact = surface.casefold()
    if exact in protected:
        return (exact,)
    lemma = lemma_table.get((exact, part_of_speech))
    return tuple(dict.fromkeys((exact, lemma))) if lemma else (exact,)

lexicon = {("returns", "VERB"): "return", ("returns", "NOUN"): "return"}
assert searchable_forms("Returns", "VERB", lexicon, set()) == ("returns", "return")
assert searchable_forms("ZX-47", "NOUN", lexicon, {"zx-47"}) == ("zx-47",)

Performance and operating cost

A dictionary lookup is O(1) average time per token and emits at most two postings here. Contextual tagging adds inference time and model storage. Duplicated postings increase index size and candidate counts, so measure retrieval quality and query latency together. The protected-term set must be versioned with the index rather than updated silently in a live query path.

Common Mistakes

  • Replacing stored source text with its lemma sequence.
  • Applying an English suffix rule to every language.
  • Reducing an exact order code because it resembles an ordinary token.
  • Claiming better search from token accuracy without checking retrieval outcomes.

Read next

Continue the workflow: Domain word senses: inventory, abstention and new meanings.

ai-data
natural-language-processing
Storage details