Chapter 32

Vision-Language Models

CLIP zero-shot, projectors and resamplers, M-RoPE, dynamic resolution, and training stages.

Vision-language models connect image representations to text representations so that a system can classify, retrieve, describe, and reason about visual input. In current LLM stacks, the vision side is usually an encoder that emits tokens, and the language side is a pretrained decoder that consumes tokens. The design question is how to align those token spaces without forgetting what either model already knows.

32.1 CLIP zero-shot classification

CLIP-style zero-shot classification embeds an image and a set of text prompts into the same space [radford2021learning]. A class such as cat is turned into prompts like "a photo of a cat"; the model picks the prompt with the largest cosine similarity to the image. With normalized embeddings i\vi and tc\vt_c, the class probabilities are

p(c∣i)=softmax⁡c(α i⊤tc),(32.1)p(c \mid \vi) = \softmax_c(\alpha\, \vi^\T \vt_c),\tag{32.1}

where α\alpha is a learned or chosen logit scale. The classifier has no task-specific head: changing the label set only changes the prompt embeddings.

Listing 32.1 Prompt similarity for CLIP-style zero-shot classification
def clip_zero_shot(image_embeddings, prompt_embeddings, logit_scale=10.0):
    """Return prompt probabilities from cosine similarities."""
    image = normalize_rows(image_embeddings)
    prompts = normalize_rows(prompt_embeddings)
    logits = logit_scale * (image @ prompts.T)
    return softmax(logits, axis=1)

Prompt wording matters because the text encoder embeds the whole phrase, not just the class name. Production systems often average several templates per class, calibrate prompts on a small validation set, or fall back to retrieval when labels are open-ended. The mathematical operation is still the same similarity matrix from Chapter 30.

Zero-shot classification is strongest when the prompts describe the visual concept in the same style as pretraining captions. A bare label may be ambiguous: "crane" could be a bird or a machine. A prompt turns the label into a short textual context, letting the text encoder place the class near the image examples it saw during contrastive training.

32.2 Projectors and learned visual queries

To feed an LLM, image tokens must have the LLM embedding width. The simplest bridge is a linear projector,

Himg=ZvisionWP+bP,(32.2)\mH_{\mathrm{img}} = \mZ_{\mathrm{vision}}\mW_P + \vb_P,\tag{32.2}

or a small MLP projector. LLaVA-style systems use this kind of bridge to connect a vision encoder to a language model, then tune on image-text instructions [liu2023visual].

Listing 32.2 Linear and MLP projectors
def linear_projector(image_tokens, weight, bias):
    """Map visual token width to the LLM embedding width."""
    return image_tokens @ weight + bias


def mlp_projector(image_tokens, w1, b1, w2, b2):
    hidden = np.tanh(image_tokens @ w1 + b1)
    return hidden @ w2 + b2

A projector preserves the number of visual tokens. A Q-Former or Perceiver-style resampler instead learns a fixed set of query tokens that cross-attend to the image tokens and return a fixed-size visual prefix [li2023blip2]. If the image encoder emits more tokens for a larger image, the LLM still receives the same number of resampled tokens.

Listing 32.3 Learned queries cross-attend to image tokens
def cross_attention(queries, context):
    """Single-head cross-attention: queries attend to context tokens."""
    scores = queries @ context.transpose(0, 2, 1) / np.sqrt(queries.shape[-1])
    weights = softmax(scores, axis=-1)
    return weights @ context, weights


def perceiver_resampler(image_tokens, learned_queries):
    """Return a fixed number of visual tokens, regardless of image token count."""
    batch = image_tokens.shape[0]
    queries = np.broadcast_to(learned_queries, (batch,) + learned_queries.shape)
    return cross_attention(queries, image_tokens)[0]

This fixed count is useful when the LLM context is the scarce resource. The cost is a possible information bottleneck: every answer must pass through the learned queries. Projector designs keep more spatial detail but spend more text-context positions.

32.3 Gated cross-attention and interleaving

Flamingo inserts cross-attention layers into a frozen language model, letting text tokens attend to visual tokens [alayrac2022flamingo]. A scalar gate controls the residual:

X′=X+tanh⁡(g) CrossAttn⁡(X,V).(32.3)\mX' = \mX + \tanh(g)\,\operatorname{CrossAttn}(\mX, \mV).\tag{32.3}

Initializing g=0g=0 makes tanh⁡(g)=0\tanh(g)=0, so the layer is exactly an identity map at the start. That protects the pretrained language model until training learns to open the gate.

Listing 32.4 Flamingo-style gated cross-attention and token interleaving
def flamingo_gated_cross_attention(text_tokens, image_tokens, gate=0.0):
    attended, _ = cross_attention(text_tokens, image_tokens)
    return text_tokens + np.tanh(gate) * attended


def interleave_image_tokens(text_tokens, image_tokens, image_at):
    before = text_tokens[:, :image_at]
    after = text_tokens[:, image_at:]
    return np.concatenate([before, image_tokens, after], axis=1)

Other VLMs place projected image tokens directly into the text sequence, for example before a question or at an <image> placeholder. Interleaving is simple once image and text tokens share the LLM width. The language model then sees a single sequence containing both modalities, while the attention mask decides which tokens can look at which earlier tokens.

32.4 M-RoPE and dynamic resolution

Text has one position axis; video and images have more. M-RoPE assigns separate temporal, height, and width position IDs to visual tokens, so a token can know which frame, row, and column it came from [wang2024qwen2vl]. For a video grid F×H×WF \times H \times W, the IDs are triples (t,h,w)(t,h,w).

Dynamic resolution keeps more tokens for large or wide images instead of forcing every image into one square. To control cost, neighboring visual tokens can be merged. With a 336×672336 \times 672 image, 14×1414 \times 14 patches produce a 24×4824 \times 48 grid: 1152 tokens. A 2×22 \times 2 merge reduces that to 288 tokens.

Listing 32.5 M-RoPE IDs and 2 by 2 token merging
def mrope_position_ids(frames, grid_h, grid_w):
    ids = []
    for t in range(frames):
        for h in range(grid_h):
            for w in range(grid_w):
                ids.append((t, h, w))
    return np.array(ids, dtype=np.int32)


def dynamic_token_counts(height, width, patch_size, merge=2):
    grid_h, grid_w = height // patch_size, width // patch_size
    if height % patch_size or width % patch_size:
        raise ValueError("height and width must be divisible by patch_size")
    if grid_h % merge or grid_w % merge:
        raise ValueError("patch grid must be divisible by merge")
    before = grid_h * grid_w
    after = (grid_h // merge) * (grid_w // merge)
    return (grid_h, grid_w), before, after


def merge_2x2_tokens(tokens, grid_hw):
    batch, _, dim = tokens.shape
    grid_h, grid_w = grid_hw
    x = tokens.reshape(batch, grid_h, grid_w, dim)
    x = x.reshape(batch, grid_h // 2, 2, grid_w // 2, 2, dim)
    return x.mean(axis=(2, 4)).reshape(batch, (grid_h // 2) * (grid_w // 2), dim)

The merge operation is usually learned or implemented as a projection after concatenating local tokens. The code averages the four tokens only to make the arithmetic visible: two spatial axes halve, so token count quarters.

32.5 Training stages

Most VLMs train in stages. Alignment pretraining teaches the bridge to connect frozen or lightly tuned vision features to text, often with caption-style data. Instruction tuning then teaches the combined model to answer visual questions, follow multimodal prompts, and produce the response format users expect. Freezing more components preserves prior knowledge and lowers cost; unfreezing more components gives better adaptation but risks forgetting.

The loss is usually ordinary language-model cross-entropy on target text. The image tokens condition the decoder; the answer tokens are predicted autoregressively. Contrastive losses may still train the vision encoder, but once a VLM is instruction-tuned, the main supervision is often "given these image tokens and prompt tokens, predict these answer tokens."

Data quality matters as much as architecture in the second stage. Captions teach recognition, but instructions teach when to answer briefly, when to refuse impossible visual claims, and how to combine visible evidence with text in the prompt. The bridge gives the LLM access to image tokens; instruction tuning teaches the behavior users expect from that access.

In practice

CLIP provides the shared image-text embedding space used for zero-shot classification and retrieval [radford2021learning]. LLaVA connects a vision encoder to an LLM with a projector and then performs visual instruction tuning [liu2023visual]. BLIP-2 uses a Q-Former to query frozen image features before handing information to a language model [li2023blip2], while Flamingo uses gated cross-attention to add visual context to a language model [alayrac2022flamingo]. Recent high-resolution VLMs use position schemes such as M-RoPE and dynamic token counts to handle images and videos with varied shapes [wang2024qwen2vl].

Key equations
p(c∣i)=softmax⁡c(α i⊤tc)p(c \mid \vi) = \softmax_c(\alpha\, \vi^\T\vt_c)
Himg=ZvisionWP+bP\mH_{\mathrm{img}} = \mZ_{\mathrm{vision}}\mW_P + \vb_P
Qlearnedattends to⁡Zvision→Q tokens\mQ_{\mathrm{learned}} \operatorname{ attends\ to } \mZ_{\mathrm{vision}} \rightarrow Q \text{ tokens}
X′=X+tanh⁡(g)CrossAttn⁡(X,V),g=0⇒X′=X\mX' = \mX + \tanh(g)\operatorname{CrossAttn}(\mX,\mV), \qquad g=0 \Rightarrow \mX'=\mX
(F,H,W)↦(t,h,w),2 by 2 merge: HW↦HW/4(F,H,W) \mapsto (t,h,w), \qquad \text{2 by 2 merge: } HW \mapsto HW/4

32.6 Teach it

The one-sentence version. A vision-language model turns images into tokens that a language model can compare with text, attend to, or read as part of its prompt.

An analogy. The vision encoder writes index cards about the image; the bridge translates the cards into the LLM’s language; the LLM answers using those cards and the user’s words.

At the board.

  1. Start with CLIP: image vector, prompt vectors, cosine similarities, softmax over labels.

  2. Show a projector that changes token width, then a resampler that changes token count.

  3. Add Flamingo’s gated cross-attention and set the gate to zero to show identity.

  4. Write M-RoPE triples (t,h,w)(t,h,w) and quarter the token count with a 2 by 2 merge.

Misconceptions to address.

  • "The LLM sees pixels." It sees embeddings or tokens produced by a vision encoder.

  • "A projector and a resampler do the same thing." One changes width; the other changes count.

  • "Dynamic resolution is free." More visual tokens spend more context and attention.

Check for understanding. Why does a zero-initialized tanh gate let us add cross-attention to a pretrained language model without changing its initial function?

32.7 Exercises

Exercise 32.1 ★ Prompt classification

Using the toy vectors in toy_zero_shot_label, compute which prompt each of two image embeddings selects. Why is no classifier head needed?

Exercise 32.2 ★★ Flamingo’s identity start

Use (32.3) to prove that a gated cross-attention block initialized with g=0g=0 is an identity map, regardless of the image tokens.

Exercise 32.3 ★★ Dynamic-resolution arithmetic

For a 336×672336 \times 672 image, 14×1414 \times 14 patches, and a 2×22 \times 2 token merge, compute the patch grid and token counts before and after merging.

Exercise 32.4 ★★★ Fixed visual prefix

Create learned queries with shape 4×d4 \times d. Show that perceiver_resampler returns four tokens for both short and long image-token sequences. Explain what information bottleneck this creates.

References

  • [alayrac2022flamingo] J. Alayrac et al. Flamingo: a Visual Language Model for Few-Shot Learning. 2022. arXiv:2204.14198

  • [li2023blip2] J. Li et al. BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models. 2023. arXiv:2301.12597

  • [liu2023visual] H. Liu et al. Visual Instruction Tuning. 2023. arXiv:2304.08485

  • [radford2021learning] A. Radford et al. Learning Transferable Visual Models From Natural Language Supervision. 2021. arXiv:2103.00020

  • [wang2024qwen2vl] P. Wang et al. Qwen2-VL: Enhancing Vision-Language Model’s Perception of the World at Any Resolution. 2024. arXiv:2409.12191