feat(phase-12/02): CLIP and contrastive vision-language pretraining

This commit is contained in:
Rohit Ghumare
2026-04-23 18:41:01 +01:00
parent 530fde07cd
commit 5e2bc2c455
5 changed files with 486 additions and 0 deletions
@@ -0,0 +1,111 @@
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 960 560" font-family="Georgia, 'Times New Roman', serif">
<defs>
<style>
.box { fill: #faf6ef; stroke: #1a1a1a; stroke-width: 1.5; }
.diag { fill: #fff1d6; stroke: #c0392b; stroke-width: 1.5; }
.neg { fill: #dfe9ff; stroke: #2c5ea9; stroke-width: 1; }
.cool { fill: #e6f4ea; stroke: #2e7d32; stroke-width: 1.5; }
.label { font-size: 13px; font-weight: 600; fill: #1a1a1a; }
.step { font-size: 12px; font-family: 'Menlo', monospace; fill: #222; }
.small { font-size: 10px; font-family: 'Menlo', monospace; fill: #555; }
.caption { font-size: 11px; fill: #555; font-style: italic; }
.title { font-size: 16px; font-weight: 700; fill: #1a1a1a; }
.head { font-size: 12px; font-weight: 700; fill: #1a1a1a; }
</style>
</defs>
<text x="480" y="24" text-anchor="middle" class="title">CLIP similarity matrix — positives on the diagonal, negatives everywhere else</text>
<rect x="40" y="50" width="440" height="480" class="box"/>
<text x="260" y="72" text-anchor="middle" class="head">similarity matrix S (N=4)</text>
<text x="260" y="88" text-anchor="middle" class="small">S[i,j] = cos(img_i, txt_j) / tau</text>
<text x="80" y="118" class="step">txt 0</text>
<text x="150" y="118" class="step">txt 1</text>
<text x="220" y="118" class="step">txt 2</text>
<text x="290" y="118" class="step">txt 3</text>
<g>
<text x="60" y="150" class="step">img 0</text>
<rect x="100" y="130" width="60" height="40" class="diag"/>
<rect x="170" y="130" width="60" height="40" class="neg"/>
<rect x="240" y="130" width="60" height="40" class="neg"/>
<rect x="310" y="130" width="60" height="40" class="neg"/>
<text x="130" y="155" text-anchor="middle" class="small">+0.82</text>
<text x="200" y="155" text-anchor="middle" class="small">-0.11</text>
<text x="270" y="155" text-anchor="middle" class="small">+0.04</text>
<text x="340" y="155" text-anchor="middle" class="small">-0.22</text>
</g>
<g>
<text x="60" y="190" class="step">img 1</text>
<rect x="100" y="170" width="60" height="40" class="neg"/>
<rect x="170" y="170" width="60" height="40" class="diag"/>
<rect x="240" y="170" width="60" height="40" class="neg"/>
<rect x="310" y="170" width="60" height="40" class="neg"/>
<text x="130" y="195" text-anchor="middle" class="small">-0.18</text>
<text x="200" y="195" text-anchor="middle" class="small">+0.77</text>
<text x="270" y="195" text-anchor="middle" class="small">+0.12</text>
<text x="340" y="195" text-anchor="middle" class="small">+0.09</text>
</g>
<g>
<text x="60" y="230" class="step">img 2</text>
<rect x="100" y="210" width="60" height="40" class="neg"/>
<rect x="170" y="210" width="60" height="40" class="neg"/>
<rect x="240" y="210" width="60" height="40" class="diag"/>
<rect x="310" y="210" width="60" height="40" class="neg"/>
<text x="130" y="235" text-anchor="middle" class="small">+0.06</text>
<text x="200" y="235" text-anchor="middle" class="small">+0.14</text>
<text x="270" y="235" text-anchor="middle" class="small">+0.79</text>
<text x="340" y="235" text-anchor="middle" class="small">-0.03</text>
</g>
<g>
<text x="60" y="270" class="step">img 3</text>
<rect x="100" y="250" width="60" height="40" class="neg"/>
<rect x="170" y="250" width="60" height="40" class="neg"/>
<rect x="240" y="250" width="60" height="40" class="neg"/>
<rect x="310" y="250" width="60" height="40" class="diag"/>
<text x="130" y="275" text-anchor="middle" class="small">-0.21</text>
<text x="200" y="275" text-anchor="middle" class="small">+0.08</text>
<text x="270" y="275" text-anchor="middle" class="small">+0.03</text>
<text x="340" y="275" text-anchor="middle" class="small">+0.84</text>
</g>
<text x="260" y="320" text-anchor="middle" class="head">InfoNCE = - sum log softmax(diag)</text>
<text x="260" y="340" text-anchor="middle" class="small">each row pushes the diagonal up, negatives down</text>
<text x="260" y="356" text-anchor="middle" class="small">symmetric: also do it column-wise</text>
<rect x="60" y="380" width="400" height="130" class="cool"/>
<text x="260" y="402" text-anchor="middle" class="head">training ingredients</text>
<text x="80" y="424" class="small">· 400M image-text pairs (CLIP), 10B+ for SigLIP 2</text>
<text x="80" y="442" class="small">· batch 32k-512k</text>
<text x="80" y="460" class="small">· learnable temperature tau (init 0.07)</text>
<text x="80" y="478" class="small">· dual encoder: ViT + small text transformer</text>
<text x="80" y="496" class="small">· normalize both embeddings before cosine</text>
<rect x="500" y="50" width="420" height="480" class="box"/>
<text x="710" y="72" text-anchor="middle" class="head">softmax vs sigmoid loss</text>
<rect x="520" y="92" width="380" height="200" class="diag"/>
<text x="540" y="114" class="step">InfoNCE (CLIP)</text>
<text x="540" y="134" class="small">per row: softmax normalizes across N</text>
<text x="540" y="150" class="small">needs full similarity matrix in sync</text>
<text x="540" y="166" class="small">distributed: all-gather every batch</text>
<text x="540" y="186" class="small">comm cost: O(world_size x batch x D)</text>
<text x="540" y="206" class="step">loss_i2t = CE(S, eye)</text>
<text x="540" y="224" class="step">loss_t2i = CE(S^T, eye)</text>
<text x="540" y="242" class="step">loss = (loss_i2t + loss_t2i) / 2</text>
<text x="540" y="268" class="small">temperature controls sharpness</text>
<text x="540" y="284" class="small">scale ceiling: 32k batch before comm dominates</text>
<rect x="520" y="300" width="380" height="210" class="cool"/>
<text x="540" y="322" class="step">Sigmoid pairwise (SigLIP)</text>
<text x="540" y="342" class="small">per pair: independent BCE</text>
<text x="540" y="358" class="small">y=1 on diagonal, y=0 off-diagonal</text>
<text x="540" y="376" class="small">loss = -y log sig(S+b) - (1-y) log sig(-S-b)</text>
<text x="540" y="394" class="small">no all-gather; local blocks only</text>
<text x="540" y="412" class="small">comm cost: O(world_size x D)</text>
<text x="540" y="432" class="step">scale ceiling: 512k+ batch feasible</text>
<text x="540" y="454" class="small">extra bias parameter b handles class imbalance</text>
<text x="540" y="472" class="small">SigLIP 2 (2025) ships with NaFlex</text>
<text x="540" y="490" class="small">+ multilingual (100+ langs)</text>
</svg>

After

Width:  |  Height:  |  Size: 6.5 KiB

@@ -0,0 +1,189 @@
"""CLIP / SigLIP contrastive loss toy — stdlib Python.
Implements InfoNCE (softmax) and sigmoid pairwise loss on a hand-constructed
similarity matrix. Also runs a tiny zero-shot-classification walkthrough using
synthetic image and text embeddings.
No numpy. No torch. The point is to see the loss math and the argmax pattern.
"""
from __future__ import annotations
import math
import random
def normalize(v: list[float]) -> list[float]:
n = math.sqrt(sum(x * x for x in v)) or 1.0
return [x / n for x in v]
def cosine(a: list[float], b: list[float]) -> float:
return sum(x * y for x, y in zip(a, b))
def similarity_matrix(images: list[list[float]],
texts: list[list[float]],
tau: float) -> list[list[float]]:
I = [normalize(v) for v in images]
T = [normalize(v) for v in texts]
N = len(I)
S = [[0.0] * N for _ in range(N)]
for i in range(N):
for j in range(N):
S[i][j] = cosine(I[i], T[j]) / tau
return S
def log_sum_exp(row: list[float]) -> float:
m = max(row)
return m + math.log(sum(math.exp(x - m) for x in row))
def infonce_loss(S: list[list[float]]) -> float:
"""Symmetric InfoNCE over rows and columns."""
N = len(S)
loss_i2t = 0.0
for i in range(N):
loss_i2t += -S[i][i] + log_sum_exp(S[i])
loss_t2i = 0.0
for j in range(N):
col = [S[i][j] for i in range(N)]
loss_t2i += -S[j][j] + log_sum_exp(col)
return (loss_i2t + loss_t2i) / (2 * N)
def sigmoid(x: float) -> float:
if x >= 0:
z = math.exp(-x)
return 1.0 / (1.0 + z)
z = math.exp(x)
return z / (1.0 + z)
def sigmoid_loss(S: list[list[float]], bias: float = 0.0) -> float:
"""SigLIP-style per-pair BCE. Positives are the diagonal."""
N = len(S)
total = 0.0
count = 0
for i in range(N):
for j in range(N):
logit = S[i][j] + bias
y = 1.0 if i == j else 0.0
p = sigmoid(logit)
eps = 1e-9
term = y * math.log(p + eps) + (1 - y) * math.log(1 - p + eps)
total += -term
count += 1
return total / count
def zero_shot_classify(image: list[float],
class_texts: dict[str, list[float]]) -> list[tuple[str, float]]:
"""Argmax cosine similarity over class prompts."""
img = normalize(image)
scores = []
for name, vec in class_texts.items():
scores.append((name, cosine(img, normalize(vec))))
scores.sort(key=lambda p: p[1], reverse=True)
return scores
def make_fake_embedding(seed: int, dim: int = 64) -> list[float]:
rng = random.Random(seed)
return [rng.gauss(0, 1) for _ in range(dim)]
def demo_infonce() -> None:
print("\nDEMO 1: InfoNCE on 4 aligned pairs")
print("-" * 60)
images = [make_fake_embedding(i) for i in range(4)]
texts = [[x + 0.05 * make_fake_embedding(i + 100)[k] for k, x in enumerate(v)]
for i, v in enumerate(images)]
for tau in (0.07, 0.1, 1.0):
S = similarity_matrix(images, texts, tau=tau)
loss = infonce_loss(S)
slip = sigmoid_loss(S)
print(f" tau={tau:4.2f} InfoNCE={loss:.4f} SigLIP={slip:.4f}")
def demo_shuffled() -> None:
print("\nDEMO 2: what happens with misaligned pairs")
print("-" * 60)
images = [make_fake_embedding(i) for i in range(6)]
texts = [make_fake_embedding(i + 500) for i in range(6)]
S = similarity_matrix(images, texts, tau=0.07)
loss = infonce_loss(S)
slip = sigmoid_loss(S)
print(f" misaligned: InfoNCE={loss:.4f} SigLIP={slip:.4f}")
aligned_imgs = [make_fake_embedding(i) for i in range(6)]
aligned_txt = [[x + 0.02 for x in v] for v in aligned_imgs]
S2 = similarity_matrix(aligned_imgs, aligned_txt, tau=0.07)
print(f" aligned : InfoNCE={infonce_loss(S2):.4f} "
f"SigLIP={sigmoid_loss(S2):.4f}")
print(" aligned loss < misaligned loss confirms the gradient signal.")
def demo_zero_shot() -> None:
print("\nDEMO 3: zero-shot classification")
print("-" * 60)
classes = {
"cat": make_fake_embedding(42),
"dog": make_fake_embedding(43),
"bird": make_fake_embedding(44),
"car": make_fake_embedding(45),
}
query_image = [c + 0.3 * make_fake_embedding(999)[i]
for i, c in enumerate(classes["dog"])]
ranked = zero_shot_classify(query_image, classes)
print(" query image (close to 'dog' prototype):")
for name, score in ranked:
print(f" {name:6s}: {score:+.4f}")
print(f" top-1: {ranked[0][0]}")
def demo_prompt_ensemble() -> None:
print("\nDEMO 4: prompt template ensemble")
print("-" * 60)
templates = [
"a photo of a {class}",
"a picture of a {class}",
"an image of a {class}",
]
class_name = "golden retriever"
ensemble_vec = [0.0] * 64
count = 0
for t in templates:
prompt = t.format(**{"class": class_name})
seed = sum(ord(c) for c in prompt)
emb = make_fake_embedding(seed)
for k in range(64):
ensemble_vec[k] += emb[k]
count += 1
ensemble_vec = [x / count for x in ensemble_vec]
print(f" ensembled {count} prompts for '{class_name}'")
print(f" first 6 dims: {[round(x, 3) for x in ensemble_vec[:6]]}")
print(" single-template: noisier; ensemble: +1-3 points on real benchmarks.")
def main() -> None:
print("=" * 60)
print("CLIP / SIGLIP CONTRASTIVE TRAINING (Phase 12, Lesson 02)")
print("=" * 60)
demo_infonce()
demo_shuffled()
demo_zero_shot()
demo_prompt_ensemble()
print("\n" + "=" * 60)
print("TAKEAWAYS")
print("-" * 60)
print(" · InfoNCE penalizes rows AND columns (symmetric)")
print(" · Lower tau -> sharper softmax -> more hard-negative pressure")
print(" · Sigmoid loss decouples pairs -> no all-gather in distributed runs")
print(" · Zero-shot = argmax cos(image, prompt) over class prompts")
if __name__ == "__main__":
main()
@@ -0,0 +1,156 @@
# CLIP and Contrastive Vision-Language Pretraining
> OpenAI's CLIP (2021) proved a single idea big enough to power the next five years: align an image encoder and a text encoder in the same vector space using only noisy web image-caption pairs and a contrastive loss. Zero supervised labels. 400M pairs. The resulting embedding space does zero-shot classification, image-text retrieval, and plugs into every 2026 VLM as its vision tower. SigLIP 2 (2025) replaced softmax with sigmoid and scaled past CLIP at lower cost. This lesson walks the math from InfoNCE to sigmoid pairwise loss and builds the training step in stdlib Python.
**Type:** Build
**Languages:** Python (stdlib, InfoNCE + sigmoid loss implementations)
**Prerequisites:** Phase 12 · 01 (ViT patches), Phase 7 (Transformers)
**Time:** ~180 minutes
## Learning Objectives
- Derive InfoNCE loss from mutual information and implement a numerically-stable vectorized version.
- Explain why sigmoid pairwise loss (SigLIP) scales to batch 32768+ without the all-gather overhead softmax demands.
- Run zero-shot ImageNet classification by constructing text templates (`a photo of a {class}`) and taking argmax over cosine similarity.
- Name the four levers CLIP / SigLIP pretraining gives you: batch size, temperature, prompt template, data quality.
## The Problem
Pre-CLIP vision was supervised. Collect labeled datasets (ImageNet: 1.2M images, 1000 classes), train a CNN, ship it. Labels are expensive, labels bias to what labelers can agree on, and labels do not transfer to new tasks without finetuning.
The image-caption web has one billion-plus loosely-labeled pairs for free. A picture of a golden retriever with alt text "my dog Max in the park" carries a supervisory signal — the text describes the image. The question: can you turn this into useful training?
CLIP's answer: treat image-caption pairs as a matching task. Given a batch of N images and N captions, learn to match each image to its own caption against N-1 distractors. The supervision is "these two things belong together; these N-1 do not." No class labels. No human annotation. Just a contrastive loss.
The resulting embedding space does more than CLIP was trained for. ImageNet zero-shot works because "a photo of a cat" embeds near pictures of cats that were never explicitly labeled cats. This is the bet that spawned every 2026 VLM.
## The Concept
### The dual encoder
CLIP has two towers:
- Image encoder `f`: ViT or ResNet, outputs a D-dim vector per image.
- Text encoder `g`: small transformer, outputs a D-dim vector per caption.
Both towers normalize their outputs to unit length. Similarity is `cos(f(x), g(y)) = f(x)^T g(y)` since both are unit-norm.
For a batch of N (image, caption) pairs, build the similarity matrix `S` of shape `(N, N)`:
```
S[i, j] = cos(f(x_i), g(y_j)) / tau
```
where `tau` is a learned temperature (CLIP initializes to 0.07; learned in log-space).
### InfoNCE loss
CLIP uses a symmetric cross-entropy over rows and columns:
```
loss_i2t = CE(S, labels=identity) # each image's positive is its own caption
loss_t2i = CE(S^T, labels=identity) # each caption's positive is its own image
loss = (loss_i2t + loss_t2i) / 2
```
This is InfoNCE. The softmax in CE forces each image to match its caption more than every other caption in the batch. The "negatives" are all other batch items. Bigger batches = more negatives = stronger signal. CLIP trained at batch 32k; scale matters.
### Temperature
`tau` controls the sharpness of the softmax. Low tau → sharp distribution, hard negative mining effect. High tau → soft, all samples contribute. CLIP learns log(1/tau), clipped to prevent collapse. SigLIP 2 fixes the initial tau and uses a learned bias instead.
### Why sigmoid scales better (SigLIP)
Softmax needs the whole similarity matrix in sync. In distributed training you must all-gather every embedding to every replica, then do the softmax. This is quadratic in world size for communication.
SigLIP replaces softmax with element-wise sigmoid: for each pair `(i, j)`, the loss is a binary classification of "are these the matching pair?" positive class labels are the diagonal, everything else is negative. The loss is:
```
L = -1/N sum over (i, j) [ y_ij log sigmoid(S[i,j]) + (1-y_ij) log sigmoid(-S[i,j]) ]
```
`y_ij = 1` if `i == j`, else 0. Each pair's loss is independent. No all-gather needed. Each GPU computes its local block and sums. SigLIP 2 scales to batch 32k-512k cheaply where CLIP would need proportionally more communication.
### Zero-shot classification
Given N class names, for each class build a text template:
```
"a photo of a {class}"
```
Embed each template with the text encoder. Embed your image with the image encoder. Argmax cosine similarity = predicted class. No training on the target classes.
Prompt templates matter. CLIP's original paper used 80 templates per class (plain, artistic, photo, painting, etc.) and averaged the embeddings. +3 ImageNet points. Modern usage typically picks one or two templates.
### Linear probes and finetuning
Zero-shot is a baseline. A linear probe (train one linear layer on top of frozen CLIP features for your target classes) beats zero-shot on in-domain tasks. Full finetuning beats linear probe on in-domain but can hurt zero-shot transfer. Three regimes with three trade-offs.
### SigLIP 2: NaFlex and dense features
SigLIP 2 (2025) adds:
- NaFlex: single model handles variable aspect ratios and resolutions.
- Better dense features for segmentation and depth estimation, targeting use as a frozen backbone in VLMs.
- Multilingual: trained on 100+ languages where CLIP was English-only.
- 1B param scale where CLIP topped out at 400M.
In 2026 open VLMs, SigLIP 2 SO400m/14 is the default vision tower. CLIP remains the default for pure image-text retrieval where the specific LAION-2B training distribution matches your query pattern.
### ALIGN, BASIC, OpenCLIP, EVA-CLIP
ALIGN (Google, 2021): same idea as CLIP, 1.8B pair scale, 90% noisy. Proved noisy data scales. OpenCLIP (LAION): open reproduction of CLIP on LAION-400M / 2B, multiple scales, the go-to open checkpoint. EVA-CLIP: initializes from masked image modeling; strong backbone for VLMs. BASIC: Google's CLIP+ALIGN hybrid. All the same family, different data and tuning.
### The zero-shot ceiling
CLIP-class models cap around 76% ImageNet zero-shot (CLIP-G, OpenCLIP-G). Beyond requires either much larger data (SigLIP 2 gets 80%+) or architecture changes (supervised heads, more parameters). The benchmark is saturating; the real value is the embedding space that downstream VLMs consume.
## Use It
`code/main.py` implements:
1. A toy dual encoder (hash-based image features, text char features) so you can see the InfoNCE shape without numpy.
2. InfoNCE loss in pure Python (numerical stability via log-sum-exp).
3. Sigmoid pairwise loss for comparison.
4. A zero-shot classification routine: compute cosine similarity against a set of text prompts, argmax for prediction.
Run it and watch the loss curve. The absolute numbers are toy; the shape matches what a real CLIP trainer emits.
## Ship It
This lesson produces `outputs/skill-clip-zero-shot.md`. Given a set of images (via path) and a list of target classes, it builds text prompts with the CLIP template, embeds both sides with a stated checkpoint (e.g., `openai/clip-vit-large-patch14`), and returns top-1 / top-5 predictions with similarity scores. The skill refuses to make claims about classes not in the prompt list.
## Exercises
1. Implement InfoNCE for a batch of 4 pairs by hand. Construct the 4x4 similarity matrix, run softmax, pick out the diagonal, compute cross-entropy. Verify your Python implementation against this hand calculation.
2. SigLIP uses a bias parameter `b` in addition to temperature: `S'[i,j] = S[i,j]/tau + b`. What role does `b` play when the batch has a large class imbalance (many more negatives than positives per row)? Read SigLIP Section 3 (arXiv:2303.15343).
3. Build a zero-shot classifier for cats vs dogs. Try two prompt templates: `a photo of a {class}` and `a picture of a {class}`. Measure accuracy on 100 test images. Does the ensemble of templates beat single?
4. Compute the communication cost of softmax InfoNCE vs sigmoid pairwise for a 512-GPU run at batch 32k. Which scales as O(N), which as O(N^2)? Cite SigLIP Section 4.
5. Read the OpenCLIP scaling-laws paper (arXiv:2212.07143, Cherti et al.). Reproduce their conclusion for data scaling from the figures: at fixed model size, what is the log-linear relationship between ImageNet zero-shot accuracy and training data size?
## Key Terms
| Term | What people say | What it actually means |
|------|----------------|------------------------|
| InfoNCE | "Contrastive loss" | Cross-entropy over a batch's similarity matrix; each item's positive is its paired item, negatives are everything else |
| Sigmoid loss | "SigLIP loss" | Per-pair binary cross-entropy; no softmax, no all-gather, scales cheaply in distributed training |
| Temperature | "tau" | Scalar that scales logits before softmax/sigmoid; controls sharpness of the distribution |
| Zero-shot | "no-finetune classification" | Use text prompts to construct class embeddings and classify by cosine similarity; no training on target classes |
| Prompt template | "a photo of a ..." | Text scaffold around a class name; affects zero-shot accuracy by 1-5 points |
| Dual encoder | "Two-tower" | One image encoder + one text encoder, outputs in shared D-dim space |
| Hard negative | "Tough distractor" | A negative similar enough to the positive that the model has to work to separate them |
| Linear probe | "Frozen + one layer" | Train only a linear classifier on top of frozen features; measures feature quality |
| NaFlex | "Native flexible resolution" | SigLIP 2 capability to ingest images at any aspect ratio and resolution without resizing |
| Temperature scaling | "log-parametrized tau" | CLIP parametrizes `log(1/tau)` so gradients behave; clips to prevent collapse to near-zero tau |
## Further Reading
- [Radford et al. — Learning Transferable Visual Models From Natural Language Supervision (arXiv:2103.00020)](https://arxiv.org/abs/2103.00020) — the CLIP paper.
- [Zhai et al. — Sigmoid Loss for Language Image Pre-Training (arXiv:2303.15343)](https://arxiv.org/abs/2303.15343) — SigLIP.
- [Tschannen et al. — SigLIP 2 (arXiv:2502.14786)](https://arxiv.org/abs/2502.14786) — multilingual + NaFlex.
- [Jia et al. — ALIGN (arXiv:2102.05918)](https://arxiv.org/abs/2102.05918) — scale with noisy web data.
- [Cherti et al. — Reproducible scaling laws for contrastive language-image learning (arXiv:2212.07143)](https://arxiv.org/abs/2212.07143) — OpenCLIP scaling laws.
@@ -0,0 +1,30 @@
---
name: clip-zero-shot
description: Run zero-shot image classification with a CLIP / SigLIP checkpoint, producing ranked predictions with similarity scores.
version: 1.0.0
phase: 12
lesson: 02
tags: [clip, siglip, zero-shot, vision-language]
---
Given a list of images (file paths or URLs) and a list of candidate class names, produce a ranked zero-shot classification using a declared CLIP or SigLIP checkpoint. The skill is pure-prediction; it does not train or finetune.
Produce:
1. Prompt construction. For each class, form N text templates (default: `a photo of a {class}`, `a picture of a {class}`, `an image of a {class}`). Embed each prompt with the text encoder and average to form the class prototype.
2. Image embedding. Embed each input image with the stated vision encoder. Normalize both sides to unit length.
3. Ranked predictions. Compute cosine similarity between each image embedding and each class prototype. Return top-1 and top-5 with scores.
4. Checkpoint metadata. Name the exact Hugging Face checkpoint used (e.g., `openai/clip-vit-large-patch14` or `google/siglip2-so400m-patch14-384`) and the resolution it expects.
5. Honesty notice. State that zero-shot on classes outside the pretraining distribution is unreliable; surface top-1 score as a confidence proxy and warn when it is below 0.2.
Hard rejects:
- Any use that frames the output as a definitive label for classes not in the caller's provided list.
- Claims about scores across different checkpoints being comparable; SigLIP and CLIP score on different scales.
- Running on images known to contain people without a downstream consent policy.
Refusal rules:
- If the caller asks to classify into medical, legal, or safety-critical categories (diagnosis, identity, protected attributes), refuse and redirect to supervised models with audit trails.
- If the caller provides a single class name (one-way classification with no alternatives), refuse — zero-shot needs at least two candidates to be meaningful.
- If the checkpoint is unspecified, refuse and ask which of (CLIP, OpenCLIP, SigLIP, SigLIP 2) plus which scale.
Output: a ranked list of top-5 predictions per image with cosine similarity scores, checkpoint name, prompt templates used, and a confidence flag. End with a "what to read next" paragraph pointing to Lesson 12.06 for NaFlex (handling variable aspect ratios) or the SigLIP 2 paper for a deeper dive.