Chapter 32
Vision-Language Models
CLIP zero-shot, projectors and resamplers, M-RoPE, dynamic resolution, and training stages.
Vision-language models connect image representations to text representations so that a system can classify, retrieve, describe, and reason about visual input. In current LLM stacks, the vision side is usually an encoder that emits tokens, and the language side is a pretrained decoder that consumes tokens. The design question is how to align those token spaces without forgetting what either model already knows.
32.1 CLIP zero-shot classification
CLIP-style zero-shot classification embeds an image and a set of text prompts into the same
space [radford2021learning]. A class such as cat is turned into prompts like "a photo of a
cat"; the model picks the prompt with the largest cosine similarity to the image. With normalized
embeddings and , the class probabilities are
where is a learned or chosen logit scale. The classifier has no task-specific head: changing the label set only changes the prompt embeddings.
def clip_zero_shot(image_embeddings, prompt_embeddings, logit_scale=10.0):
"""Return prompt probabilities from cosine similarities."""
image = normalize_rows(image_embeddings)
prompts = normalize_rows(prompt_embeddings)
logits = logit_scale * (image @ prompts.T)
return softmax(logits, axis=1)
Prompt wording matters because the text encoder embeds the whole phrase, not just the class name. Production systems often average several templates per class, calibrate prompts on a small validation set, or fall back to retrieval when labels are open-ended. The mathematical operation is still the same similarity matrix from Chapter 30.
Zero-shot classification is strongest when the prompts describe the visual concept in the same style as pretraining captions. A bare label may be ambiguous: "crane" could be a bird or a machine. A prompt turns the label into a short textual context, letting the text encoder place the class near the image examples it saw during contrastive training.
32.2 Projectors and learned visual queries
To feed an LLM, image tokens must have the LLM embedding width. The simplest bridge is a linear projector,
or a small MLP projector. LLaVA-style systems use this kind of bridge to connect a vision encoder to a language model, then tune on image-text instructions [liu2023visual].
def linear_projector(image_tokens, weight, bias):
"""Map visual token width to the LLM embedding width."""
return image_tokens @ weight + bias
def mlp_projector(image_tokens, w1, b1, w2, b2):
hidden = np.tanh(image_tokens @ w1 + b1)
return hidden @ w2 + b2
A projector preserves the number of visual tokens. A Q-Former or Perceiver-style resampler instead learns a fixed set of query tokens that cross-attend to the image tokens and return a fixed-size visual prefix [li2023blip2]. If the image encoder emits more tokens for a larger image, the LLM still receives the same number of resampled tokens.
def cross_attention(queries, context):
"""Single-head cross-attention: queries attend to context tokens."""
scores = queries @ context.transpose(0, 2, 1) / np.sqrt(queries.shape[-1])
weights = softmax(scores, axis=-1)
return weights @ context, weights
def perceiver_resampler(image_tokens, learned_queries):
"""Return a fixed number of visual tokens, regardless of image token count."""
batch = image_tokens.shape[0]
queries = np.broadcast_to(learned_queries, (batch,) + learned_queries.shape)
return cross_attention(queries, image_tokens)[0]
This fixed count is useful when the LLM context is the scarce resource. The cost is a possible information bottleneck: every answer must pass through the learned queries. Projector designs keep more spatial detail but spend more text-context positions.
32.3 Gated cross-attention and interleaving
Flamingo inserts cross-attention layers into a frozen language model, letting text tokens attend to visual tokens [alayrac2022flamingo]. A scalar gate controls the residual:
Initializing makes , so the layer is exactly an identity map at the start. That protects the pretrained language model until training learns to open the gate.
def flamingo_gated_cross_attention(text_tokens, image_tokens, gate=0.0):
attended, _ = cross_attention(text_tokens, image_tokens)
return text_tokens + np.tanh(gate) * attended
def interleave_image_tokens(text_tokens, image_tokens, image_at):
before = text_tokens[:, :image_at]
after = text_tokens[:, image_at:]
return np.concatenate([before, image_tokens, after], axis=1)
Other VLMs place projected image tokens directly into the text sequence, for example before a
question or at an <image> placeholder. Interleaving is simple once image and text tokens share
the LLM width. The language model then sees a single sequence containing both modalities, while
the attention mask decides which tokens can look at which earlier tokens.
32.4 M-RoPE and dynamic resolution
Text has one position axis; video and images have more. M-RoPE assigns separate temporal, height, and width position IDs to visual tokens, so a token can know which frame, row, and column it came from [wang2024qwen2vl]. For a video grid , the IDs are triples .
Dynamic resolution keeps more tokens for large or wide images instead of forcing every image into one square. To control cost, neighboring visual tokens can be merged. With a image, patches produce a grid: 1152 tokens. A merge reduces that to 288 tokens.
def mrope_position_ids(frames, grid_h, grid_w):
ids = []
for t in range(frames):
for h in range(grid_h):
for w in range(grid_w):
ids.append((t, h, w))
return np.array(ids, dtype=np.int32)
def dynamic_token_counts(height, width, patch_size, merge=2):
grid_h, grid_w = height // patch_size, width // patch_size
if height % patch_size or width % patch_size:
raise ValueError("height and width must be divisible by patch_size")
if grid_h % merge or grid_w % merge:
raise ValueError("patch grid must be divisible by merge")
before = grid_h * grid_w
after = (grid_h // merge) * (grid_w // merge)
return (grid_h, grid_w), before, after
def merge_2x2_tokens(tokens, grid_hw):
batch, _, dim = tokens.shape
grid_h, grid_w = grid_hw
x = tokens.reshape(batch, grid_h, grid_w, dim)
x = x.reshape(batch, grid_h // 2, 2, grid_w // 2, 2, dim)
return x.mean(axis=(2, 4)).reshape(batch, (grid_h // 2) * (grid_w // 2), dim)
The merge operation is usually learned or implemented as a projection after concatenating local tokens. The code averages the four tokens only to make the arithmetic visible: two spatial axes halve, so token count quarters.
32.5 Training stages
Most VLMs train in stages. Alignment pretraining teaches the bridge to connect frozen or lightly tuned vision features to text, often with caption-style data. Instruction tuning then teaches the combined model to answer visual questions, follow multimodal prompts, and produce the response format users expect. Freezing more components preserves prior knowledge and lowers cost; unfreezing more components gives better adaptation but risks forgetting.
The loss is usually ordinary language-model cross-entropy on target text. The image tokens condition the decoder; the answer tokens are predicted autoregressively. Contrastive losses may still train the vision encoder, but once a VLM is instruction-tuned, the main supervision is often "given these image tokens and prompt tokens, predict these answer tokens."
Data quality matters as much as architecture in the second stage. Captions teach recognition, but instructions teach when to answer briefly, when to refuse impossible visual claims, and how to combine visible evidence with text in the prompt. The bridge gives the LLM access to image tokens; instruction tuning teaches the behavior users expect from that access.
|
In practice
|
CLIP provides the shared image-text embedding space used for zero-shot classification and retrieval [radford2021learning]. LLaVA connects a vision encoder to an LLM with a projector and then performs visual instruction tuning [liu2023visual]. BLIP-2 uses a Q-Former to query frozen image features before handing information to a language model [li2023blip2], while Flamingo uses gated cross-attention to add visual context to a language model [alayrac2022flamingo]. Recent high-resolution VLMs use position schemes such as M-RoPE and dynamic token counts to handle images and videos with varied shapes [wang2024qwen2vl]. |
32.6 Teach it
The one-sentence version. A vision-language model turns images into tokens that a language model can compare with text, attend to, or read as part of its prompt.
An analogy. The vision encoder writes index cards about the image; the bridge translates the cards into the LLM’s language; the LLM answers using those cards and the user’s words.
At the board.
-
Start with CLIP: image vector, prompt vectors, cosine similarities, softmax over labels.
-
Show a projector that changes token width, then a resampler that changes token count.
-
Add Flamingo’s gated cross-attention and set the gate to zero to show identity.
-
Write M-RoPE triples and quarter the token count with a 2 by 2 merge.
Misconceptions to address.
-
"The LLM sees pixels." It sees embeddings or tokens produced by a vision encoder.
-
"A projector and a resampler do the same thing." One changes width; the other changes count.
-
"Dynamic resolution is free." More visual tokens spend more context and attention.
Check for understanding. Why does a zero-initialized tanh gate let us add cross-attention to a pretrained language model without changing its initial function?
32.7 Exercises
Using the toy vectors in toy_zero_shot_label, compute which prompt each of two image
embeddings selects. Why is no classifier head needed?
Use (32.3) to prove that a gated cross-attention block initialized with is an identity map, regardless of the image tokens.
For a image, patches, and a token merge, compute the patch grid and token counts before and after merging.
Create learned queries with shape . Show that perceiver_resampler returns
four tokens for both short and long image-token sequences. Explain what information bottleneck
this creates.
References
-
[alayrac2022flamingo] J. Alayrac et al. Flamingo: a Visual Language Model for Few-Shot Learning. 2022. arXiv:2204.14198
-
[li2023blip2] J. Li et al. BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models. 2023. arXiv:2301.12597
-
[liu2023visual] H. Liu et al. Visual Instruction Tuning. 2023. arXiv:2304.08485
-
[radford2021learning] A. Radford et al. Learning Transferable Visual Models From Natural Language Supervision. 2021. arXiv:2103.00020
-
[wang2024qwen2vl] P. Wang et al. Qwen2-VL: Enhancing Vision-Language Model’s Perception of the World at Any Resolution. 2024. arXiv:2409.12191