Meta’s Llama 4 is mindblowing… but did it cheat?

FireshipAbout 4 min readMay 27, 2025Watch original
THE SUMMARYAI-generated

Key Concepts:

  • Llama 4 (Maverick, Scout, Behemoth): Meta's open-weight, natively multimodal mixture of experts large language model family.
  • Context Window: The amount of text a language model can consider at once.
  • LM Arena Leaderboard: A platform where language models are ranked based on head-to-head human preference comparisons.
  • Open Model vs. Open-Source: Open models are free to use but not necessarily with fully accessible and modifiable source code.
  • Multimodal: Ability to process different types of data, such as text, images, and video.
  • AI Agent: An AI system designed to perform specific tasks autonomously.
  • Vibe Coding: A term implying coding based on intuition or feeling rather than rigorous methodology.
  • Augment Code: An AI agent designed for large-scale codebases.

1. Llama 4 Release and Initial Impressions

  • Meta released Llama 4, an open-weight, natively multimodal mixture of experts large language model family. It includes three models: Maverick, Scout, and Behemoth.
  • Scout boasts an unprecedented 10 million token context window.
  • Initially, Llama 4 topped the LM Arena leaderboard, outperforming proprietary models except for Gemini 2.5 Pro.
  • However, the model on LM Arena was allegedly a fine-tuned version ("impostor") optimized for human preference, not the actual open-weight Llama 4.
  • Ella Marina criticized Meta's interpretation of their policy, stating that Llama 4 "did not match what we expect from model providers" and "wasn't passing the vibe check."

2. Shopify's AI-First Strategy

  • A leaked internal memo from Shopify's CEO revealed an "AI-first strategy."
  • Teams must justify why they cannot use AI for tasks before requesting more resources or headcount.
  • Employees are expected to learn AI, implying potential job displacement for those who don't adapt.
  • The memo highlights the economic incentives for companies to adopt AI, replacing human labor.
  • The memo's transparency is appreciated, especially in the context of open models like Llama 4.

3. Llama 4 Model Details and Performance

  • Llama 4 is natively multimodal, capable of understanding image and video inputs.
  • Scout has a 10 million token context window, while Maverick has a 1 million token context window. Behemoth is still in training.
  • Despite the large context window, real-world performance on large codebases is reportedly underwhelming.
  • Memory requirements for utilizing the 10 million token context window are substantial.
  • The internet community has expressed disappointment with Llama 4's overall performance.

4. Benchmarking Controversy

  • Llama 4's strong benchmark performance led to accusations of intentional training on testing data.
  • Meta has denied these accusations.
  • Despite being considered a "flop" by some, Llama 4 is still a valuable resource as an open model (though not fully open-source).

5. Augment Code: AI Agent for Codebases

  • Augment Code is presented as an AI agent designed for large-scale codebases.
  • It understands a team's entire codebase, enabling it to perform tasks like migrations and testing.
  • Augment Code integrates with tools like VS Code, GitHub, and Vim.
  • It learns and fine-tunes itself based on a team's unique coding style.
  • A developer plan is available with unlimited usage.

6. Technical Terms and Concepts

  • Open-weight: Refers to models where the weights (parameters) are publicly available.
  • Natively Multimodal: The model is designed from the ground up to handle multiple data types.
  • Mixture of Experts: An architecture where multiple specialized models are combined.
  • Context Window: The maximum sequence length a model can process at once.
  • Token: A unit of text used by the model (e.g., a word or subword).
  • Fine-tuning: Adapting a pre-trained model to a specific task or dataset.

7. Logical Connections

  • The video connects the release of Llama 4 with the broader trend of AI adoption in the workplace, exemplified by Shopify's memo.
  • It contrasts the initial hype surrounding Llama 4's benchmark performance with the reported real-world limitations.
  • The video transitions from discussing general-purpose language models to highlighting specialized AI agents like Augment Code.

8. Synthesis/Conclusion

Llama 4's release generated significant buzz due to its multimodal capabilities and large context window. However, questions about its real-world performance and potential benchmark manipulation have tempered expectations. The Shopify memo underscores the growing pressure on employees to adapt to AI-driven workflows. While open models like Llama 4 offer valuable resources, specialized AI agents like Augment Code are emerging to address specific industry needs, such as codebase management. The video highlights the ongoing evolution of AI and its impact on both technology and the job market.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.