Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Project: audit incident text index contracts

Last updated: 5 Oct 202630 min read
project
IntermediateBy AITrove Editorial

An incident console needs four different text lookups. One request asks for documents containing every supplied exact term. Another asks for consecutive tokens in a fixed order. A third asks whether a character fragment occurs anywhere in stored source text. The final request asks whether one vocabulary term is present in a compact sorted dictionary. Build the reference results by scanning token lists, testing adjacent slices, applying source-text substring checks, and checking membership in a plain set. Then compare those results with sorted postings, positional postings, a trigram filter, and front-coded blocks. Do not treat these as interchangeable queries: document membership loses position, character fragments ignore token boundaries, and vocabulary membership does not identify any document.

Acceptance trace

Use incident IDs 103, 218, 347, and 492. The document-term query valve AND pressure returns 103 and 347; repeated pressure in one record does not duplicate its ID. For the token sequences in the phrase lesson, pump pressure returns 103, 218, and 347, while pressure pressure returns only 347. For the raw text records, pressure returns 103, 218, and 492 after full-fragment verification. A three-term first block reconstructs pressure, pressurize, and pressurized; pressurizer remains absent. Try an empty query, an absent term, repeated phrase tokens, and a two-character fragment.

Expected output

Output
and=[103, 347] phrase=[103, 218, 347] substring=[103, 218, 492]
lexicon-hit=true lexicon-miss=false

Boundary and cost review

State the tokenization and case rules once, then apply them before both indexing and querying. Use unique IDs for documents and distinguish a term's occurrence frequency from its document frequency. Measure the extra stored positions for phrase queries and the retained raw text needed to verify trigram candidates. Compare front-coded payload sizes with ordinary strings in the actual runtime rather than assuming a byte saving from a conceptual encoding. All four examples build snapshots: a changed document or vocabulary entry requires rebuilding here. A production incremental index would need update, deletion, and publication rules not supplied by these lessons.

Common Mistakes

  • Do not duplicate an ID in one term's posting list.
  • Do not answer an exact phrase with document-level AND alone.
  • Do not return trigram candidates without complete substring verification.
  • Do not interpret a decoded shared prefix as a term hit.

Connected lessons

data structures
projects
Storage details