Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Embedding index migration: bind vectors, queries and source revisions

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

An embedding migration changes the retrieval contract, even when both indexes accept vectors of the same length.

Define the contract before rebuilding

A repair-manual search service holds service bulletins as chunks. Its existing index uses an older embedding model; a new model changes vector meaning. Record model digest, tokenizer and preprocessing revision, vector dimension, distance metric, normalization policy, chunking revision, source snapshot and metadata filters as one index contract. A dimension match alone cannot prove compatibility. The query encoder must share the document encoder revision for the index receiving the request. Generative release binding extends that contract to the response model.

Rebuild from stable sources

Take a source snapshot with document IDs and revision timestamps. Generate chunks and embeddings under the new contract, then write them to a separate collection. Preserve source IDs and chunk lineage so deleted bulletins can be removed from both collections. Count expected and indexed records by source; compare hashes of source revisions at completion. A backfill that finishes while source documents keep changing can still be stale. Change capture can supply updates, but the consumer still needs a replayable checkpoint.

Carry live updates through the handover

Keep the old collection serving. During backfill, either dual-write source changes to both collections or append them to a durable change log and replay them into the new collection before cutover. Put a monotonic source revision in each update. If revision 82 arrives before revision 47, the older update must not replace it. Record deletes as tombstones rather than quietly leaving expired chunks searchable. Training-serving parity teaches the analogous issue for tabular features.

Budget storage and prove completeness

An approximate vector index can use far more memory than raw float storage because graph and metadata structures add overhead. Plan for a period with two full indexes and a growing replay log. Publish per-source counts, delete coverage, source-revision lag and rejected vectors. Hold promotion when the new collection lacks an eligible bulletin even if sample queries look good. Shadow evaluation then measures relevance and latency on a frozen query set before an alias changes.

Implementation

python
def index_contract_matches(request, index):
    fields = ("encoder_digest", "dimension", "metric", "normalization",
              "chunk_revision")
    mismatch = [field for field in fields if request.get(field) != index.get(field)]
    return "admit" if not mismatch else "reject:" + ",".join(mismatch)

index = {"encoder_digest": "manual-embed-r47", "dimension": 384,
         "metric": "cosine", "normalization": "unit", "chunk_revision": "r8"}
query = dict(index)
assert index_contract_matches(query, index) == "admit"
assert index_contract_matches({**query, "encoder_digest": "manual-embed-r46"},
                              index) == "reject:encoder_digest"

Performance and operating cost

The contract check is O(k) time and O(k) temporary space for k fields. Backfill uses O(n × d) embedding work for n chunks of dimension d; two indexes temporarily double base vector storage before index overhead. Query admission is cheap. Embedding generation and index construction dominate migration cost.

Common Mistakes

  • Assuming equal vector dimensions mean equal semantic spaces.
  • Replacing an index in place while old query encoders still serve.
  • Ignoring deletes that occurred during backfill.
  • Comparing corpus counts without source revisions or eligible filters.

Read next

ai-data
mlops
Storage details