Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Text embeddings: pair labels, hard negatives and versioned vectors

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

A semantic search model learns the retrieval task only when its positive and negative pairs reflect what the user needs, and the stored vectors match the active encoder.

Define relevance at the request level

Suppose an engineer searches “payment callback received twice” and needs the runbook that explains duplicate delivery. A positive pair is a query with a passage that resolves that request, not merely one containing the same product word. A negative pair is a reviewed non-answer. Record the query, source document revision, passage boundaries, reviewer decision and reason. Multiple passages may be relevant. Avoid assigning every unclicked result a negative label: it might have been unseen, below the fold or withheld by access controls. Corpus identity and label policy applies to search judgments too.

Mine difficult negatives carefully

Random unrelated passages teach very little after the first training stage. Add near-miss negatives: a runbook for a similar callback, a stale version or a passage that mentions the same error code without resolution. Check them with reviewers because a supposedly negative passage may answer the query. Keep documents and near duplicates in one train or test group; otherwise a copied paragraph in both splits inflates retrieval quality. Grouped temporal validation is the split boundary.

Treat vectors as a derived index

The stored vector is a function of passage text, chunk policy, normalization, encoder weights and pooling rule. If any changes, rebuild or segregate the index. Mixing old and new embeddings can produce rankings that look plausible but are meaningless. Record the source revision and an encoder digest beside every vector. Keep an index alias that can move atomically after validation, and retain the prior version for rollback. Exact identifiers still need lexical matching; dense similarity alone can miss a rare incident code.

Measure retrieval before generation

Evaluate recall at a fixed candidate depth, mean reciprocal rank for first useful hit, and failures by query type. Use a time-split audit with queries written after the candidate documents were available. Report “no relevant document” cases separately; forcing a plausible hit is a product error. These measurements feed hybrid ranking and the runbook project, while retrieval systems cover the broader architecture.

Implementation

python
from hashlib import sha256

def vector_record(document_id, revision, passage, encoder_digest, embedding):
    if not passage.strip() or not embedding:
        raise ValueError("passage and embedding are required")
    passage_digest = sha256(passage.encode("utf-8")).hexdigest()
    return {"document_id": document_id, "revision": revision,
            "passage_digest": passage_digest, "encoder_digest": encoder_digest,
            "embedding": tuple(float(value) for value in embedding)}

indexed = vector_record("runbook-47", "rev-3", "Retry the callback once.",
                        "encoder-2026-10", [0.31, -0.42, 0.18])
assert indexed["document_id"] == "runbook-47"

Performance and operating cost

Encoding n passages costs n model calls and stores O(n·d) vector numbers for dimension d, before index overhead. A brute-force query scores O(n·d); an approximate index trades memory and recall for lower query latency. Index rebuilds are batch operations and must finish before the alias moves. Track encoder cost, candidate recall and exact-code misses together, because cheaper vectors can be more expensive if they increase failed searches.

Common Mistakes

  • Treating unclicked results as verified negatives.
  • Mixing embeddings from two encoder versions in one searchable index.
  • Splitting near-duplicate documents between train and audit sets.
  • Reporting only average similarity instead of relevance at a fixed candidate depth.

Read next

Continue the workflow: Hybrid text ranking with access filters and reranking.

Continue the workflow: Project: stage support-ticket duplicate groups for review.

Continue the workflow: Passage revisions: chunk migrations, deletion and audit.

ai-data
natural-language-processing
Storage details