THE SUMMARYAI-generated
Multimodal Foundation Models: A Deep Dive
Key Concepts:
- Foundation Models: Pre-trained models on a wide variety of tasks, adaptable to specific tasks with minimal data.
- Multimodal Learning: Combining different modalities (e.g., image and text) to create more robust and generalizable models.
- Self-Supervised Learning: Training models on unlabeled data using pretext tasks (e.g., contrastive learning).
- Contrastive Learning: Learning representations by pulling together similar examples and pushing apart dissimilar examples.
- Zero-Shot Learning: Performing tasks without any task-specific training data.
- Few-Shot Learning: Performing tasks with only a few examples of task-specific training data.
- Vision Language Models (VLMs): Models that combine vision and language understanding capabilities.
- Chaining: Combining multiple foundation models to perform complex tasks.
- Hallucination: When a model generates outputs that are not grounded in reality or the input data.
1. Introduction to Foundation Models
- Traditional approach: Collect dataset, train specialized model, evaluate.
- Foundation models: Pre-train on diverse tasks, adapt to specific tasks with minimal data.
- Example: GPT, pre-trained on Common Crawl data, fine-tuned for math, reasoning, trivia.
- Benefits: Minimal data required for adaptation, sometimes even zero-shot learning is possible.
- Types of foundation models:
- Language: ELMo, BERT, GPT, T5
- Image Classification: CLIP, CoCa
- Multimodal: Combining language and vision models
- Output Generation: Models that output text, masks, or images
- Chaining: Combining multiple foundation models
2. Characteristics of Foundation Models
- Robustness and Generality: Applicable to many different tasks.
- Large Scale: Large number of parameters, large amounts of training data.
- Self-Supervised Training: Trained with self-supervised objectives.
3. Image Classification with CLIP
- Building a Foundation Model for Image Classification:
- Leverage self-supervised learning (e.g., SimCLR).
- SimCLR: Contrast against dissimilar images, pull together representations of transformed images.
- Goal: General representations that can classify new concepts (e.g., sketches).
- Incorporating Text with CLIP:
- Embed text descriptions alongside image representations.
- Example: "A cute fluffy cat" close to cat images, "My favorite dog is a golden retriever" close to golden retriever images.
- CLIP Training Objective:
- Image encoder and text encoder.
- Contrastive loss: Pull together representations of similar images and text, push apart dissimilar representations.
- Symmetric loss: Image closest to its text, text closest to its image.
- CLIP Training Data:
- Image-text pairs from the internet.
- OpenAI trained CLIP on a massive dataset of internet data.
- CLIP Adaptation:
- Pre-train image encoder.
- Add a linear layer on top for specific tasks (image classification, detection, segmentation).
- CLIP achieved significant performance improvements with this linear addition.
- Zero-Shot CLIP:
- Goal: Use CLIP directly out of the box without retraining.
- Clever trick: Use the text encoder to guide the model to generalize.
- Process:
- Embed categories (e.g., plane, dog, bird) in the text space.
- Embed the input image using the image encoder.
- Find the closest neighbor in the text space.
- Essentially building a one-nearest neighbor algorithm.
- Using phrases instead of single words improves performance (e.g., "a photo of a plane").
- Averaging multiple phrases further improves performance.
- CLIP Generalization:
- CLIP generalizes well to new datasets (e.g., ObjectNet) that contain out-of-domain objects.
- Reasons for generalization:
- Text contains more structural information (shape, color).
- Scale of data (billions of image-text pairs).
- CLIP generalizes to sketches and adversarial datasets.
4. Factors Contributing to CLIP's Success
- Model Size: Transformer architecture with 307 million parameters.
- Data Size: 400 million image-text pairs from the internet.
5. CoCa: An Improvement over CLIP
- CoCa adds a decoder to the CLIP model.
- Decoder captions the image using image features from the encoder.
- Hypothesis: Captioning requires learning richer information.
- CoCa outperforms CLIP across ImageNet variants (10% boost).
- Turning point: Foundation models surpass supervised learning models for image encoders.
6. Advantages of CLIP
- Easy to train (simple contrastive learning).
- Fast inference (retrieval on embedded data).
- Open vocabulary (any text description).
- Amenable to chaining with other models.
7. Limitations of CLIP
- Compositionality: Struggles with understanding relationships between objects (e.g., mug in grass vs. grass in mug).
- Batch Size Dependence: Performance depends on large batch sizes (32,000).
- Hard Negatives: Training with hard negatives can unlearn semantics.
- Lack of Grounding: Missing information about object locations and relationships.
- Data Filtering: Even with billions of images, data filtering is crucial.
8. Vision and Language Models (VLMs)
- Motivation: Leverage the next token prediction of language models for image models.
- LLaVA: One of the first popular multimodal language models.
- ViLBERT (2019): Introduced the idea of combining image and language models.
- LLaVA Architecture:
- Feed image tokens into a language model along with historical context.
- Use CLIP image encoder to extract tokens.
- Extract features from the penultimate layer of the CLIP encoder.
- Pass features through a linear layer to convert them into something the LLM can understand.
- Flamingo:
- Follows LLaVA's setup but innovates on feature fusion.
- Feeds vision encoder features into every layer of the LLM.
- Adds a GATED X cross-attention module to every LLM layer.
- Adds a perceiver sampler to downsample image representations.
- Flamingo Training:
- Concatenation of multiple images and descriptions.
- Masking scheme to ensure descriptions only look at the corresponding image features.
- Flamingo Applications:
- Multi-turn dialogue about images.
- In-context learning.
- Classification.
- OCR and math.
- Performance: Flamingo achieves significant improvements across many benchmarks.
9. The Rise of API Models and the Open Source Gap
- Companies like OpenAI and Google release high-performing API models (GPT-4V, Gemini).
- Open source models lag behind in performance.
- Problem: Lack of understanding in the research community on how to build performant VLMs.
- Distilled models are not truly open because they rely on the existence of the original model.
10. Molmo: Closing the Open Source Gap
- Completely open source VLM (open weights, data, and code).
- Achieves competitive performance with API models.
- User study shows Molmo has a similar Elo rating to GPT-40.
- 7-billion parameter model can run on a single GPU.
- Key innovation: Grounding decision-making in the pixels themselves.
- Trained on hand-curated data with dense descriptions of image content.
- Elicitation studies to identify missing information from the internet.
- Model architecture is similar to existing models, but the data is the key difference.
- Molmo can perform tasks like pointing to specific objects in an image.
- Can be chained with other models like SAM 2 for segmentation.
11. Generalizing to Any Output Space: Segment Anything Model (SAM)
- Goal: Build a segmentation model that's a foundation model for all kinds of segmentation tasks.
- Generalize to any category and output masks for user-specified objects.
- SAM Architecture:
- Image encoder (e.g., CLIP encoder).
- Prompt encoder (encodes text, points, bounding boxes).
- Lightweight decoder (outputs a mask).
- SAM handles ambiguity by outputting three segmentation masks at different levels of granularity.
- Data: SAM authors significantly grew the amount of segmentation data.
- In-the-loop process: Human annotators refine model-generated segments.
12. Chaining: Combining Foundation Models
- Combine the strengths of different models to enable new capabilities.
- Example: Use GPT to generate descriptions for categories that CLIP hasn't seen.
- VisProg: Generate a program that uses multiple models to answer a question.
- Example: Use object detection models to count people in two images and then add the results together.
- Chaining can be used for image editing tasks.
- Can be thought of as an agent that decides which models to use for a given task.
13. Conclusion
- Foundation models enable generalization to many different downstream applications.
- Multimodal learning combines different modalities to create more robust models.
- Data quality and density are crucial for achieving high performance.
- Chaining allows combining multiple models to perform complex tasks.
- Hallucinations are still a challenge, but techniques like grounding and verification can help mitigate them.
- Future directions include building models that can create new tools and dynamically adapt to new tasks.
AI summaries can miss context or contain errors. Check important details against the original video.