Build a versioned candidate index for a learning feed, then test full-corpus recall, eligibility, new-content freshness and request latency.
Project: retrieve learning articles with two neural towers
Define the reader event ledger
Store article impressions, clicks and completed lessons with reader ID, event time, surface and policy version. Build a reader state from events strictly before the candidate decision, then pair it with a completed article that was eligible and exposed. Split by reader and time for a final evaluation. Keep new readers and new articles as named slices. The interaction contract explains why a click and a completed lesson should not have identical weight.
Train with deliberate negatives
The snippet runs one synthetic matched-pair update and asserts that each positive article appears once in the batch. Real training must deduplicate item IDs or support multiple positives, and may add reviewed hard negatives only from eligible exposed items. Compare against a topic-match baseline and a popularity baseline with identical catalog and time boundaries. The embedding lesson gives the loss and false-negative caveat.
Build a coupled index
Export item vectors alongside encoder and feature-schema fingerprints. Build a candidate index from published articles, maintain tombstones and check incremental additions. At serving, encode the reader with the matching query tower, retrieve more than the final count, filter eligibility and return a bounded list. Compare exact top-K and indexed top-K on fixed queries to isolate index error. The index lesson defines the replay oracle.
Evaluate the feed, not the training batch
Report full-corpus recall at candidate count, eligible yield, catalog coverage and recall by new reader, new article and topic. Separately audit withdrawn-item leaks and the delay before a new article appears. Measure p95 query encoding, index search, filtering and total request time on the target catalog size. A candidate system that retrieves only popular old pages can look strong on interaction-weighted recall while failing discovery.
Release with a clear fallback
Package both towers, article vectors, index version, eligibility snapshot rule and rollback pointer. Run a shadow replay first, compare candidate sets and investigate blocked items. Gate on full-corpus recall, freshness, zero restricted-item exposure and request latency. Keep a simple eligible-topic baseline ready if the index is stale or unavailable; never relax eligibility to fill the feed.
Implementation
import torch
from torch import nn
from torch.nn import functional as functional
torch.manual_seed(47)
reader_context = torch.rand(4, 5)
article_metadata = torch.rand(4, 7)
positive_article_ids = [47, 58, 69, 81]
assert len(positive_article_ids) == len(set(positive_article_ids))
query_tower = nn.Linear(5, 8)
item_tower = nn.Linear(7, 8)
optimizer = torch.optim.AdamW(list(query_tower.parameters()) +
list(item_tower.parameters()), lr=0.0007)
query_vectors = functional.normalize(query_tower(reader_context), dim=1)
item_vectors = functional.normalize(item_tower(article_metadata), dim=1)
candidate_logits = query_vectors @ item_vectors.T / 0.17
loss = functional.cross_entropy(candidate_logits, torch.arange(4))
optimizer.zero_grad(set_to_none=True)
loss.backward()
optimizer.step()
assert candidate_logits.shape == (4, 4) and torch.isfinite(loss)Performance and operating cost
The in-batch B-by-B score matrix costs O(B squared D) operations for embedding width D and O(B squared) memory, on top of tower computation. An index for M published items stores O(MD) vector values plus search metadata. Exact full-catalog evaluation is more expensive than a small training-batch check; sample queries but keep the corpus complete and frozen. The code proves one synthetic update, not recommendation quality or index readiness.
Common Mistakes
- Do not treat a single batch loss as a feed-quality metric.
- Do not publish an item-vector index without encoder and schema fingerprints.
- Do not bypass eligibility when the index returns too few candidates.
Read next
- Two-tower embeddings and in-batch negative contracts
- Retrieval index versions, eligibility filters and candidate parity
- Offline recommendation evaluation: replay catalog and learner state in time
- Project: build a guarded next-lesson recommender
- Inference contracts: preserve preprocessing and measure tail latency
