How AI connects text and images

3Blue1BrownAbout 2 min readAug 22, 2025Watch original
THE SUMMARYAI-generated

Key Concepts:

  • Diffusion Models: AI models that generate images and videos by reversing a diffusion process analogous to Brownian motion.
  • Brownian Motion: The random movement of particles in a fluid due to collisions.
  • CLIP (Contrastive Language-Image Pre-training): A model architecture composed of a text encoder and an image encoder that learns to associate images and their captions in a shared embedding space.
  • Embedding Space: A high-dimensional vector space where images and text are represented as vectors, with similar concepts being closer together.
  • Text Encoder: A component of CLIP that transforms text into a vector representation.
  • Image Encoder: A component of CLIP that transforms images into a vector representation.

1. Diffusion Models and Brownian Motion:

  • AI systems excel at converting text prompts into videos using diffusion models.
  • Diffusion models operate based on a process remarkably equivalent to Brownian motion.
  • Brownian motion is the random movement of particles, and diffusion models reverse this process in high-dimensional space.

2. CLIP Model Architecture:

  • In February 2021, OpenAI released CLIP, a model architecture with two components: a text encoder and an image encoder.
  • Both encoders output vectors of length 512.
  • The core idea is that vectors representing an image and its corresponding caption should be similar in the embedding space.

3. Mathematical Operations in Embedding Space:

  • Example: Taking two pictures, one of the speaker without a hat and one with a hat.
  • Passing both images through the CLIP image model yields two vectors in the embedding space.
  • Subtracting the "no hat" vector from the "hat" vector results in a new vector.
  • This new vector represents the mathematical difference between the two concepts.

4. Textual Interpretation of Vector Differences:

  • The new vector (hat - no hat) can be used to search for corresponding text.
  • By passing various words through the text encoder, the model can identify the closest textual match.
  • In the example, the top-ranked match is "hat," followed by "cap" and "helmet."

5. Learned Geometry of Embedding Space:

  • CLIP's embedding space allows for mathematical operations on pure ideas or concepts in images and text.
  • The model has learned a geometry where semantic relationships are reflected in vector relationships.

6. Conclusion:

  • CLIP demonstrates the ability to perform mathematical operations on abstract concepts within its embedding space.
  • This capability allows the model to understand and manipulate the relationships between images and text in a meaningful way.
  • The learned geometry of the embedding space is crucial for the model's ability to associate images and text effectively.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.