Meta's Llama 4 Models: Scout, Maverick, and Behemoth - A Disappointing Review
Key Concepts:
- Mixture of Experts (MoE)
- Parameters (Total vs. Active)
- Context Window
- Inference Speed
- Benchmarks
- Censorship
- Reasoning Model
- Omni Model
- Open Weights
- Fine-tuning
- Multimodal
Model Overview
Meta has released three new large language models (LLMs):
- Scout: 109 billion total parameters, 17 billion active parameters, 16 experts, 10 million token context window.
- Llama 4 Maverick: 400 billion parameters, 17 billion active parameters, 128 experts.
- Behemoth: 2 trillion parameters, 288 billion active parameters, 16 experts.
All three models utilize a Mixture of Experts (MoE) architecture, where similar vectors are grouped into one expert, theoretically making each expert specialized in a specific task (e.g., math, coding).
Technical Limitations
Despite the MoE architecture, none of the models are locally usable on consumer-grade GPUs. Even with only 17 billion active parameters, the entire model (e.g., 109 billion for Scout) needs to be loaded into memory, which exceeds the capacity of typical GPUs. The MoE architecture primarily benefits inference speed.
Benchmark Performance
The benchmark results are disappointing:
- Scout: Performs similarly to Gemma 3, despite having approximately five times the capacity.
- Maverick: Performs similarly to 2.0 Flash and GPT-4o.
- Behemoth: Reportedly beats 3.7 Sonnet in benchmarks.
The models are less censored than previous iterations. Scout and Maverick are available for download on Hugging Face, but Behemoth is only mentioned in a blog post and is not available for use or download. A reasoning model and an omni (multimodal) version are expected soon.
Accessibility
The models are accessible through the Meta AI platform, Groq, Together AI, Hyperbolic, and Open Router. Open Router offers a free API for both Scout and Maverick.
Testing and Results
The video author tested Scout and Maverick on a series of questions, revealing significant shortcomings, especially in coding tasks.
- Passes: Both models passed simple questions and a confetti button coding task.
- Fails: Both models failed complex reasoning and pattern recognition questions. Maverick passed a question that Scout struggled with.
- Coding Catastrophes: The models performed poorly on more complex coding tasks:
- Playable Synth Keyboard: The generated code was "atrocious" and non-functional.
- Butterfly SVG Generation: The output was unacceptable for a 400 billion parameter model; smaller models like Fi 414b can generate better SVGs.
- Spinning Hexagon: The code was broken, with the ball sliding off the hexagon.
- Game of Life: Maverick's implementation was functional, but Scout's did not work at all.
Overall Assessment
The models are considered "pretty bad by today's standards," even considering their size. The author expresses surprise at the poor performance, especially given the anticipation surrounding Llama 4. He speculates that the poor performance is why Meta delayed the launch.
Recommendations
The author does not recommend using these models, particularly for coding. He suggests that smaller models like Flash are superior. While the open weights might make them suitable for fine-tuning, the base performance is lacking, especially in coding. The 10 million token context window could be beneficial for document understanding and multimodal applications.
Usage with Clients
Users can integrate the models with clients like Klein or Rue Code by upgrading the client and setting the provider to Open Router or Groq, selecting the Llama 4 free endpoint.
Conclusion
Meta's new Llama 4 models, particularly Scout and Maverick, are a disappointment, especially in coding tasks. The author recommends waiting for the reasoning and omni models. While the large context window might be useful for specific applications, the overall performance does not justify their use compared to existing alternatives.
AI summaries can miss context or contain errors. Check important details against the original video.