Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Two-tower embeddings and in-batch negative contracts

Last updated: 7 Oct 20265 min read
tutorial
AdvancedBy AITrove Editorial

Separate a query encoder from an item encoder so candidate vectors can be indexed, while keeping positives, accidental duplicates and exposure bias visible during training.

Give each tower one job

A learning feed can encode a reader’s recent topics in a query tower and encode an article’s topic, level and text in an item tower. Both emit vectors with the same dimension; their dot product supplies a retrieval score. Item vectors can be prepared ahead of time, leaving only the query tower and nearest-neighbor search on the request path. Keep the feature schema, normalization and index version together. Candidate eligibility still filters items a user cannot or should not see.

Construct positives from eligible exposure

A completed lesson is a useful positive signal, but an unseen lesson is not automatically a disliked negative. Record what was actually shown, when, where and under which ranking policy. For held-out evaluation, split by time and reader before building training pairs. A query representation that includes clicks after the target event leaks the answer. Exposure bias explains why observed interactions are not a complete preference map.

Handle false negatives within the batch

A common training matrix pairs each query with its own positive item and treats other items in that mini-batch as negatives. If two readers completed the same article, that article appears twice and the off-diagonal copy is a false negative. Build a mask by item ID or use a multi-positive objective. The code validates unique item IDs for a deliberately simple diagonal loss; a real batch builder must enforce that invariant or change the objective. Hard negatives need the same eligibility and time checks.

Choose score scale and embedding norm

Dot-product logits depend on vector magnitude; normalized vectors yield cosine-like scores, while a temperature controls softmax sharpness during training. Neither choice alone fixes biased negatives. Monitor embedding norms, score distribution and query groups. A model can learn popularity alone if item frequency overwhelms the query signal. Compare against a topical or popularity baseline, especially for new readers and new articles.

Evaluate full-corpus retrieval

A batch classification loss tests only the current mini-batch. Report recall at a chosen candidate count against the entire eligible article corpus or a stable evaluation snapshot, then inspect recall by topic, reader history and publication age. The nearest-neighbor index can lose some exact neighbors; measure that gap separately from model quality. The index lesson makes that parity check explicit.

Implementation

python
import torch
from torch import nn
from torch.nn import functional as functional

torch.manual_seed(47)
reader_features = torch.rand(4, 5)
article_features = torch.rand(4, 7)
positive_article_ids = [47, 58, 69, 81]
assert len(set(positive_article_ids)) == len(positive_article_ids)
reader_tower = nn.Linear(5, 8)
article_tower = nn.Linear(7, 8)
reader_vectors = functional.normalize(reader_tower(reader_features), dim=1)
article_vectors = functional.normalize(article_tower(article_features), dim=1)
pair_scores = reader_vectors @ article_vectors.T / 0.17
matched_positions = torch.arange(len(reader_features))
loss = functional.cross_entropy(pair_scores, matched_positions)
loss.backward()
assert pair_scores.shape == (4, 4) and torch.isfinite(loss)

Performance and operating cost

For a batch of B queries and B candidate items with embedding width D, the score matrix costs O(B squared D) arithmetic and O(B squared) memory, which can dominate large-batch training. Precomputing M item vectors costs O(MD) index storage. Query encoding plus approximate search is cheaper than scanning all M items at serve time, but index build, filtering and freshness add separate costs. This synthetic update checks tensor geometry only and assumes unique positives.

Common Mistakes

  • Do not treat unexposed articles as confirmed dislikes.
  • Do not allow duplicate positive item IDs to become false in-batch negatives.
  • Do not use mini-batch accuracy as full-corpus retrieval recall.

Read next

ai-data
deep-learning
Storage details