Stanford CS336 Language Modeling from Scratch | Spring 2026 | Lecture 17: Alignment - Multimodality
By Stanford Online
Key Concepts
- Multimodality: The ability of a model to process and generate information across different types of data (text, images, audio, video).
- Omnimodel: A theoretical model capable of ingesting and outputting any combination of modalities.
- CLIP (Contrastive Language-Image Pre-training): A foundational model that aligns images and text in a shared embedding space using contrastive learning.
- SigLIP (Sigmoid Loss for Language Image Pre-training): An improvement over CLIP that uses a sigmoid loss function, decoupling batch size from the loss and improving training efficiency.
- VLM (Vision Language Model): Models that combine a vision encoder with a Large Language Model (LLM) to perform visual reasoning.
- AnyRes: A methodology for handling arbitrary image resolutions by breaking images into patches/crops to maintain fine-grained detail.
- VQ-VAE (Vector Quantized Variational Autoencoder): A technique to map images into discrete tokens, allowing them to be processed by standard language models.
- Multimodal RoPE (Rotary Positional Embeddings): A technique to encode spatial and temporal coordinates (height, width, time) into transformer inputs.
1. Foundations of Multimodal Models
The lecture emphasizes that while Transformers were designed for text (discrete tokens), they are the most effective architecture for all modalities at scale. The core challenge is converting non-text data (images, video, audio) into "tokens" (discrete or continuous) that a Transformer can process.
- CLIP (2021):
- Objective: Align image and text embeddings using a dot-product similarity. It treats the task as a multiclass classification problem across a batch of $N$ image-text pairs.
- Architecture: Uses a Vision Transformer (ViT) as the image encoder and a GPT-2 style Transformer as the text encoder.
- Key Insight: By training on massive, noisy web-scraped data, CLIP learns robust semantic representations that outperform models trained on curated datasets like ImageNet.
- SigLIP:
- Innovation: Replaces the multiclass softmax with a binary sigmoid loss for each image-text pair.
- Benefit: It is more computationally efficient and allows for smaller batch sizes without degrading performance, addressing the scaling limitations of CLIP.
2. Vision Language Models (VLMs): LLaVA and Qwen
These models follow a "stitching" approach: taking a pre-trained vision encoder and a pre-trained LLM and connecting them via an adapter.
- LLaVA (Large Language-and-Vision Assistant):
- Process: Uses a two-stage training approach: (1) Alignment phase (freezing encoders, training the projection matrix $W$) and (2) Fine-tuning phase (training the LLM on instruction-following data).
- Data: Relies heavily on distilling GPT-4 to generate complex reasoning conversations based on image captions.
- Qwen-VL/Qwen2-VL/Qwen3-VL:
- Evolution: Qwen models evolved to handle dynamic resolutions (AnyRes) and long-context video understanding.
- Technical Improvements: Introduced Multimodal RoPE to handle 3D spatial-temporal coordinates and explicit video timestamps to improve temporal reasoning.
- Adapter Sophistication: Moved from simple linear projections to complex cross-attention and deep fusion into the residual stream of the LLM.
3. The "Chameleon" Approach: Discrete Tokens
The Chameleon model (Meta) attempts to treat images as discrete tokens (via VQ-VAE) so that the entire model is a standard language model.
- Pros: Elegant, unified architecture; no need for specialized adapters.
- Cons: Training instability due to high entropy in image tokens; loss of fine-grained information (e.g., OCR performance suffers).
4. Methodologies and Frameworks
- Data Curation: Modern VLMs rely on "post-training" where high-quality, task-specific data (VQA, OCR, chart reasoning) is synthesized using stronger models like GPT-4.
- Training Stages: Most state-of-the-art VLMs use a multi-stage pipeline:
- Alignment: Connecting the vision encoder to the LLM.
- Knowledge/Task Tuning: Training on high-quality, domain-specific datasets.
- Instruction Tuning: Final refinement for conversational capabilities.
- Handling Resolution: The "AnyRes" strategy is critical—downsampling the full image for global context while cropping patches for high-resolution detail.
5. Key Arguments and Insights
- Semantic vs. Fine-Grained: CLIP-style encoders are excellent for high-level semantics (classification) but struggle with fine-grained tasks like OCR. Diffusion models are preferred for generation because they handle high-frequency details better.
- System Constraints: Multimodal training is significantly more complex than text-only training. Data loading is a bottleneck, and balancing the information density between modalities (e.g., video vs. text) requires careful weighting and normalization.
- Transfer Learning: The lecturer notes that models trained on single-image tasks (like OCR or diagrams) often show emergent capabilities in multi-image or video reasoning, suggesting that the underlying Transformer architecture effectively generalizes spatial relationships.
Conclusion
The field is moving toward "natively multimodal" models. While current open-source approaches (LLaVA, Qwen) rely on stitching vision encoders to LLMs, the future likely involves continuous encoders for understanding and diffusion-based heads for generation. The primary takeaway is that data curation and task-specific distillation are currently more important for performance than architectural breakthroughs.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
