An incident console needs four different text lookups. One request asks for documents containing every supplied exact term. Another asks for consecutive tokens in a fixed order. A third asks whether a character fragment occurs anywhere in stored source text. The final request asks whether one vocabulary term is present in a compact sorted dictionary. Build the reference results by scanning token lists, testing adjacent slices, applying source-text substring checks, and checking membership in a plain set. Then compare those results with sorted postings, positional postings, a trigram filter, and front-coded blocks. Do not treat these as interchangeable queries: document membership loses position, character fragments ignore token boundaries, and vocabulary membership does not identify any document.
Project: audit incident text index contracts
Acceptance trace
Use incident IDs 103, 218, 347, and 492. The document-term query valve AND pressure returns 103 and 347; repeated pressure in one record does not duplicate its ID. For the token sequences in the phrase lesson, pump pressure returns 103, 218, and 347, while pressure pressure returns only 347. For the raw text records, pressure returns 103, 218, and 492 after full-fragment verification. A three-term first block reconstructs pressure, pressurize, and pressurized; pressurizer remains absent. Try an empty query, an absent term, repeated phrase tokens, and a two-character fragment.
Expected output
and=[103, 347] phrase=[103, 218, 347] substring=[103, 218, 492]
lexicon-hit=true lexicon-miss=falseBoundary and cost review
State the tokenization and case rules once, then apply them before both indexing and querying. Use unique IDs for documents and distinguish a term's occurrence frequency from its document frequency. Measure the extra stored positions for phrase queries and the retained raw text needed to verify trigram candidates. Compare front-coded payload sizes with ordinary strings in the actual runtime rather than assuming a byte saving from a conceptual encoding. All four examples build snapshots: a changed document or vocabulary entry requires rebuilding here. A production incremental index would need update, deletion, and publication rules not supplied by these lessons.
Common Mistakes
- Do not duplicate an ID in one term's posting list.
- Do not answer an exact phrase with document-level AND alone.
- Do not return trigram candidates without complete substring verification.
- Do not interpret a decoded shared prefix as a term hit.
