Stanford CS25: V5 I Transformers for Video Generation, Andrew Brown of Meta

Unknown AuthorAbout 6 min readJul 4, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Text-to-Video Generation: Creating videos from textual descriptions.
  • Transformers: A neural network architecture that excels at processing sequential data.
  • Scaling Laws: Empirical relationships between model size, compute, and performance.
  • Temporal Autoencoder (TAE): A variational autoencoder used for spatial-temporal video compression.
  • Flow Matching: A generative modeling technique, a simpler generalization of diffusion modeling, used to learn the distribution of video data.
  • Llama 3: Meta's large language model architecture, used as the base for Movie Gen.
  • Cross-Attention: A mechanism for incorporating text conditioning into the video generation process.
  • Adaptive Layer Norm (AdaLN): A technique for conditioning the model on the time step in flow matching.
  • Data Filtration: The process of cleaning and curating a high-quality video dataset for training.

1. Introduction and Background

  • Andrew Brown, a research scientist at Meta's Gen AI team, discusses the use of transformers for video generation, focusing on the Movie Gen model.
  • The talk highlights the rapid progress in video generation, contrasting state-of-the-art results from October 2024 with those from September 2022, emphasizing the significant advancements achieved in a short period.
  • The core argument is that scaling data, compute, and model parameters for a simple transformer architecture, combined with flow matching, enables high-quality video generation.

2. Historical Context and Movie Gen Overview

  • Two milestone events in video generation are identified: the adoption of diffusion modeling in 2022 and the shift towards transformer-based architectures in 2024.
  • Before 2024, specialized architectures like CNNs and U-Nets were prevalent, but the field has since moved towards unified transformer architectures for efficiency and scalability.
  • Movie Gen is presented as a family of foundation models capable of generating high-quality 1080p HD videos with synchronized audio. The talk primarily focuses on the text-to-video model.
  • Movie Gen is a 30 billion parameter model trained on approximately 100 million videos and 1 billion images.

3. Model Architecture: Representation

  • The representation of data is crucial for generative modeling, specifically how to represent videos (X) for the model to learn P(X).
  • Text data is highly compressed and semantically rich, while media data is continuous, raw, and contains significant redundancy.
  • Modeling pixels directly is computationally expensive, scaling quadratically with resolution, making it impractical for high-resolution videos.
  • Prior work uses compressed latent representations learned via VAEs or VQ-VAEs trained offline to alleviate the computational burden.
  • Movie Gen employs a Temporal Autoencoder (TAE) for spatial-temporal video compression, achieving 8x compression in height, width, and time.
  • The TAE consists of an encoder and a decoder. The encoder compresses the video into a latent representation, which is then decoded back to pixel space.
  • The largest video modeled in the paper is 768x768 pixels, 16 seconds, 16 fps, compressed from 150 million tokens (pixels) to 73,000 tokens using the TAE.

4. Model Architecture: Learning Objective

  • Unlike text generation, which often uses autoregression and next token prediction, media generation commonly employs diffusion modeling or flow matching.
  • Flow matching is presented as a simpler generalization of diffusion, offering more robust training and efficient probability paths.
  • Flow Matching Process:
    1. Sample a training data point X1 (image of a cat).
    2. Sample a time step t (float between 0 and 1) and noise z from a Gaussian distribution.
    3. Construct a training sample Xt, a noised version of the image, using linear interpolation.
    4. Train the model to predict the velocity, the direction to move Xt back towards X1.
  • The learning objective is the mean squared error between the model's predicted velocity and the ground truth velocity, conditioned on the text prompt (P) and time step (t).
  • Inference involves sampling from a Gaussian distribution and using an ordinary differential equation (ODE) solver to iteratively denoise the sample back to the data distribution.

5. Model Architecture: Transformer

  • Movie Gen utilizes the Llama 3 architecture, a dense, fully connected decoder-only language model, as the base transformer.
  • The video is encoded with the TAE, flattened into a sequence of tokens, and fed into the Llama 3 architecture.
  • The Llama 3 architecture is used as a randomly initialized architecture, not a pre-trained Llama model.
  • Modifications to Llama 3:
    1. Text Conditioning: Incorporated using cross-attention layers between self-attention and feed-forward networks in the transformer block. Pre-trained frozen text models (UL2, Meta CLIP, and ByT5) are used to generate text representations.
    2. Time Step Conditioning: Implemented using adaptive layer norm (AdaLN) blocks.
    3. Bidirectional Attention: Causal masking is removed to allow every video token to attend to every other token. Multi-head attention is used instead of grouped query attention.

6. Data and Training Recipe

  • Data is crucial for training large language models, requiring internet-scale, clean data.
  • Significant resources are dedicated to data curation, often with data teams outnumbering modeling teams.
  • The model was trained on approximately 100 million videos, obtained through a detailed pipeline with handcrafted and model-based filters.
  • Data Filtration Steps:
    • Visual filtering (size, scene changes, aesthetics).
    • Motion filtering (slow, janky motion).
    • Content filtering (deduplication, resampling to balance concept distribution).
    • Automatic caption generation using Llama 3.
  • Multi-Stage Training Recipe:
    1. 256p text-to-image (T2I) stage for rapid initial training.
    2. Pre-training stage with joint text-to-image and text-to-video generation, progressively increasing resolution from 256p to 768p.
    3. Text-to-video post-training stage (SFT) on a small set of high-quality videos.

7. Results and Applications

  • The model generates high-quality videos, demonstrating reasoning about objects, motion, and physics.
  • Examples include a sloth with pink sunglasses on a donut float, showcasing generalization to novel concepts.
  • Movie Gen Edit allows for precise video editing based on text instructions.
  • A personalization model enables conditioning on an image of a person to generate videos featuring that individual.
  • Movie Gen Audio generates synchronized audio for generated videos.
  • Human evaluation studies show that Movie Gen outperformed prior work in 2024 across various metrics (motion quality, text alignment, visual quality).

8. Scaling Laws

  • Scaling laws for transformers appear to be modality-independent.
  • The Llama 3 scaling law for text serves as a reasonable predictor for model size and compute for video generation.
  • Scaling laws are used to estimate the optimal model size for a given compute budget.

9. Future Directions

  • Movie Gen did not solve video generation, and the model struggles with complex motions from complex prompts.
  • Future Research Areas:
    1. Scaling everything more (data, compute, model parameters).
    2. Incorporating reasoning capabilities (chain of thought, self-correction).
    3. Native multimodal generation (integrating video generation with other modalities).

10. Conclusion

  • The Movie Gen project demonstrates that scaling data, compute, and model parameters for a simple transformer architecture, combined with flow matching, enables high-quality video generation.
  • The architecture is modality-independent, and the scaling laws for transformers appear to be consistent across modalities.
  • Future research directions include scaling, reasoning, and native multimodal generation.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.