Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Index freshness: publish complete revisions and remove stale chunks

Last updated: 5 Oct 20265 min read
tutorial
IntermediateBy AITrove Editorial

Index updates need a versioned build, deletion handling and an atomic reader switch so answers never mix policy generations.

Build a candidate generation

When the support handbook changes, parse and index the new revision under a candidate generation. Keep the current generation serving while the candidate is incomplete. Validate document counts, active revisions, permissions and representative retrieval questions, then switch the reader pointer. Backfill publication] uses the same complete-before-visible idea.

Propagate deletion

A revoked policy can remain in an embedding index, lexical index or answer cache. Track document-to-chunk lineage so removal reaches every derived store. If a deletion cannot be completed promptly, deny access to the document ID at query time while reindexing catches up. Do not rely on an old vector disappearing merely because its source file was removed.

Choose freshness guarantees

Record ingestion lag, last successful index generation and maximum acceptable age. An urgent policy correction may require a synchronous block on old content; a routine edit may tolerate a short lag. Communicate when the answer feature is stale or unavailable rather than silently serving a superseded rule. No-answer behavior] can be safer during a gap.

Drill a partial failure

Fail one document during candidate indexing. The current generation should remain visible, the candidate should not publish, and an operator should see the rejected document and reason. Then revoke a permission and confirm query-time filtering blocks the old chunk before the next full rebuild.

Implementation

python
def publish_index(candidate, current_pointer, required_documents):
    if (set(candidate["document_ids"]) != set(required_documents)
            or len(candidate["document_ids"]) != len(set(candidate["document_ids"]))):
        raise ValueError("index generation is incomplete")
    if candidate["failed_chunks"] or not candidate["permission_checks_passed"]:
        raise ValueError("index generation failed checks")
    previous_generation = current_pointer.read()
    if not current_pointer.compare_and_swap(previous_generation, candidate["generation_id"]):
        raise RuntimeError("index pointer changed during publish")
    return previous_generation

Performance and operating cost

A full rebuild reads O(C) corpus content and creates O(K) chunks, embeddings and index entries. A pointer switch is cheap; storing two generations temporarily increases storage and build cost.

Common Mistakes

  • Do not expose a partial candidate index.
  • Do not forget answer caches when deleting content.
  • Do not wait for full reindexing to revoke access.

Read next

Continue the workflow: Exact versus approximate nearest-neighbor audit.

ai-data
retrieval-ai
Storage details