Chapter 30
Contrastive & Metric Learning
Contrastive, triplet, InfoNCE, CLIP, and SigLIP losses, and embedding retrieval.
Contrastive learning turns related objects into nearby vectors and unrelated objects into separated vectors. In LLM systems, those vectors power retrieval, ranking, deduplication, and image-text alignment before a generator ever sees a token. The core pattern is simple: score all pairs in a batch, make matched pairs win, and backpropagate through the scores.
30.1 Embeddings and cosine scores
An embedding model maps an input to a vector . We compare two rows by the cosine of the angle between them, the same geometry as Chapter 2:
Normalizing rows first makes every dot product a cosine. That matters because the model cannot win by increasing vector norms; it must rotate matched examples together and mismatched examples apart.
def normalize_rows(x, eps=1e-12):
"""Return rows scaled to unit length."""
x = np.asarray(x)
norm = np.linalg.norm(x, axis=1, keepdims=True)
return x / np.maximum(norm, eps)
def cosine_similarity(a, b):
"""All pairwise cosine similarities between row embeddings."""
return normalize_rows(a) @ normalize_rows(b).T
A retrieval system builds a matrix of query-document similarities and sorts each row. The diagonal is the target in paired data: image matches caption , question matches passage , and so on. The rest of this chapter changes only the loss placed on that matrix.
Cosine similarity also separates training from serving. During training the model sees a whole batch and can compare every item with every other item. During serving a single query is embedded once, then compared against stored normalized vectors by dot product. The loss must therefore teach the geometry directly: nearby means interchangeable for the task, not merely large in norm or convenient for a classifier head.
30.2 Margins: pairs and triplets
A margin loss states how far is far enough. Let be cosine distance and for a matching pair, otherwise. The pairwise contrastive loss is
The positive term pulls matches to distance zero. The negative term pushes only impostors that are inside the margin ; once , they stop contributing. A triplet loss compares one anchor, one positive, and one negative:
The useful negative is rarely a random one. Hard-negative mining picks the nonmatching item with the largest similarity in the current batch, so the model spends updates on mistakes it actually makes.
Margins make these losses local. A positive pair keeps pulling until it is nearly identical under the chosen distance, but a negative pair is forgotten once it crosses the margin. This is efficient when labels say only "same" or "different"; it is also brittle when two negatives are semantically close, because the loss has no way to say "farther than the positive, but not too far." Triplets fix part of that by comparing a negative to the positive for the same anchor.
def pairwise_contrastive_loss(a, b, same, margin=0.5):
"""Mean y d^2 + (1-y) max(0, margin-d)^2 for cosine distance d."""
same = np.asarray(same, dtype=np.float64)
d = cosine_distance(a, b)
pull = same * d**2
push = (1.0 - same) * np.maximum(0.0, margin - d) ** 2
return float(np.mean(pull + push))
def hard_negative_indices(anchor, candidates):
"""For paired rows, choose the nonmatching candidate with largest cosine."""
scores = cosine_similarity(anchor, candidates)
scores = scores.copy()
np.fill_diagonal(scores, -np.inf)
return np.argmax(scores, axis=1)
def triplet_loss_hard_negative(anchor, positive, margin=0.2):
"""Mean max(0, d(a,p)-d(a,n)+margin), mining n from the batch."""
negatives = positive[hard_negative_indices(anchor, positive)]
d_pos = cosine_distance(anchor, positive)
d_neg = cosine_distance(anchor, negatives)
return float(np.mean(np.maximum(0.0, d_pos - d_neg + margin)))
30.3 InfoNCE and its gradient
InfoNCE turns each row of similarities into a classification problem. With temperature , logits are , and the correct class for row is column :
Let . Differentiate the log-softmax: the derivative of contributes at the target, and the derivative of contributes . Averaging over rows gives
A smaller sharpens the softmax and also scales the gradient. The tests check this backward pass against central finite differences.
The sign is the whole algorithm. For the diagonal, is negative unless the model is already certain, so gradient descent raises the matched score. For an off-diagonal entry, is positive, so gradient descent lowers that impostor score. Hard negatives arise automatically: an off-diagonal with high probability gets the largest push down.
def info_nce_loss_and_grad(similarity, temperature=0.1):
"""Cross-entropy over rows, with correct pair on the diagonal."""
n = similarity.shape[0]
probabilities = softmax_rows(similarity / temperature)
loss = -np.log(np.diag(probabilities)).mean()
grad = probabilities.copy()
grad[np.arange(n), np.arange(n)] -= 1.0
grad /= n * temperature
return float(loss), grad
30.4 CLIP and SigLIP losses
CLIP uses both retrieval directions. If predicts the matching text from an image and predicts the matching image from text, then
The symmetric loss makes every image compete for texts and every text compete for images [radford2021learning]. SigLIP removes the row softmax. It labels diagonal pairs and off-diagonal pairs , then applies a binary logistic loss with a learnable bias :
The bias shifts the threshold between positive and negative pairs, which is useful because a batch has many more negatives than positives.
The two objectives disagree about how examples compete. In CLIP, each row and column is a probability distribution; making one negative lower raises the relative probability of all the others. In SigLIP, every pair is judged independently by a sigmoid, so the loss can be evaluated or sharded without a global softmax over the whole batch. The learnable bias absorbs the class imbalance created by many off-diagonal negatives.
def clip_loss_and_grad(similarity, temperature=0.1):
"""Symmetric CLIP loss: row retrieval plus column retrieval."""
row_loss, row_grad = info_nce_loss_and_grad(similarity, temperature)
col_loss, col_grad_t = info_nce_loss_and_grad(similarity.T, temperature)
return 0.5 * (row_loss + col_loss), 0.5 * (row_grad + col_grad_t.T)
def siglip_loss_and_grad(similarity, bias=0.0):
"""Pairwise sigmoid loss with +1 labels on the diagonal and -1 elsewhere."""
n = similarity.shape[0]
labels = -np.ones_like(similarity)
labels[np.arange(n), np.arange(n)] = 1.0
logits = labels * (similarity + bias)
loss = np.logaddexp(0.0, -logits).mean()
grad_logits = -labels / (1.0 + np.exp(logits)) / similarity.size
grad_bias = float(np.sum(grad_logits))
return float(loss), grad_logits, grad_bias
30.5 Retrieval on synthetic pairs
A minimal two-tower retriever has one encoder for each side. Here both encoders are only linear maps, trained by the symmetric CLIP loss on synthetic paired observations. The evaluation is recall@k: the fraction of queries whose true partner appears among the top scores.
def make_synthetic_pairs(n=48, latent_dim=4, image_dim=7, text_dim=6, seed=0):
rng = np.random.default_rng(seed)
z = rng.standard_normal((n, latent_dim)).astype(np.float32)
image_map = rng.standard_normal((latent_dim, image_dim)).astype(np.float32)
text_map = rng.standard_normal((latent_dim, text_dim)).astype(np.float32)
images = z @ image_map + 0.05 * rng.standard_normal((n, image_dim))
texts = z @ text_map + 0.05 * rng.standard_normal((n, text_dim))
return images.astype(np.float32), texts.astype(np.float32)
def retrieval_recall_at_k(image_embeddings, text_embeddings, k=1):
scores = cosine_similarity(image_embeddings, text_embeddings)
topk = np.argsort(-scores, axis=1)[:, :k]
target = np.arange(scores.shape[0])[:, None]
return float(np.mean(np.any(topk == target, axis=1)))
def train_linear_pair(images, texts, embed_dim=4, steps=250, lr=0.8, seed=1):
rng = np.random.default_rng(seed)
w_image = (0.1 * rng.standard_normal((images.shape[1], embed_dim)))
w_text = (0.1 * rng.standard_normal((texts.shape[1], embed_dim)))
w_image, w_text = w_image.astype(np.float32), w_text.astype(np.float32)
for _ in range(steps):
_, grad_image, grad_text = linear_pair_loss_and_grad(images, texts,
w_image, w_text)
w_image -= lr * grad_image.astype(np.float32)
w_text -= lr * grad_text.astype(np.float32)
return images @ w_image, texts @ w_text
The point is not the architecture; it is the pressure from the loss. A batch supplies positives on the diagonal and negatives everywhere else. After training, recall@1 on the seeded synthetic problem is high, while an untrained slice of the raw features is near chance.
This tiny example uses full-batch training only to keep the code readable. Real systems sample many batches, refresh mined negatives, and evaluate both directions of retrieval. The same recall@k calculation still applies: a query succeeds if its paired item appears before enough wrong items in the sorted similarity row.
|
In practice
|
Modern embedding systems use these losses at very large batch sizes because every extra paired example adds many in-batch negatives. CLIP-style image-text pretraining remains the standard way to align vision and language encoders [radford2021learning]. CPC introduced InfoNCE as a contrastive predictive objective [oord2018], FaceNet popularized triplet loss for metric embeddings [schroff2015facenet], and SigLIP replaces the softmax competition with pairwise sigmoid terms [zhai2023sigmoid]. Production retrievers usually add mined negatives from an index, not only from the current minibatch. |
30.6 Teach it
The one-sentence version. Contrastive learning trains embeddings by making the correct pair score higher than the alternatives.
An analogy. Arrange name tags and faces on a table: pull the matching tag toward each face, and push away the most confusing wrong tag.
At the board.
-
Draw two normalized vectors and write cosine as a dot product on the unit sphere.
-
For a positive/negative pair, draw the margin: negatives outside it are ignored.
-
Turn a row of pair scores into a softmax classifier; the target is the diagonal.
-
Add the column direction to get CLIP, then replace the softmax with independent sigmoid decisions to get SigLIP.
Misconceptions to address.
-
"Only positives matter." The negatives define what counts as close.
-
"Hard negatives are always mislabeled." They are often valid confusions; use labels and filters when mining them.
-
"Temperature is just a constant." It changes both probabilities and gradient scale.
Check for understanding. If a batch doubles in size, how many off-diagonal image-text comparisons does the CLIP similarity matrix contain, and why does that help retrieval training?
30.7 Exercises
For two logits whose unscaled gap is , compute the target softmax probability at and . What changes besides the probability?
For and all similarities and bias equal to zero, compute the derivative of (30.7) with respect to . Explain the sign.
Use make_synthetic_pairs, train_linear_pair, and retrieval_recall_at_k to train a small
linear image-text retriever. Report recall@1 and recall@5 before and after training.
References
-
[oord2018] A. van den Oord, Y. Li, and O. Vinyals. Representation learning with contrastive predictive coding. 2018. arXiv:1807.03748
-
[radford2021learning] A. Radford et al. Learning Transferable Visual Models From Natural Language Supervision. 2021. arXiv:2103.00020
-
[schroff2015facenet] F. Schroff, D. Kalenichenko, and J. Philbin. FaceNet: A Unified Embedding for Face Recognition and Clustering. 2015. arXiv:1503.03832
-
[zhai2023sigmoid] X. Zhai et al. Sigmoid Loss for Language Image Pre-Training. 2023. arXiv:2303.15343