THE SUMMARYAI-generated
Grok 4: Quick Thoughts and Analysis
Key Concepts:
- Grok 4: XAI's new large language model.
- Benchmarks: Standardized tests to evaluate model performance.
- Humanities Last Exam: A challenging exam used to assess reasoning abilities.
- ARC AGI 2: A benchmark focusing on abstract reasoning and generalization.
- RL Compute: Reinforcement learning used to fine-tune the model.
- Tool Usage: The ability of the model to access and use external tools.
- Multi-Agentic System: Using multiple AI agents with different tools to solve problems.
- Artificial Analysis Intelligence Index: An independent benchmark for evaluating AI models.
- Reasoning Tokens: Tokens generated by the model during its reasoning process.
Benchmarks and Performance
- Grok 4 achieves state-of-the-art results on almost all benchmarks.
- Specifically highlighted is the Humanities Last Exam, where it scores up to 50%.
- Direct comparison with other models is difficult because they lack tool access.
- On ARC AGI 2, Grok 4 achieves 16%, outperforming previous models like Claude Opus 4 (8%).
- Artificial Analysis Intelligence Index gives Grok 4 a score of 73, surpassing other models.
How Grok 4 Achieved Its Performance
- Key component: 10x more RL compute compared to Grok 3 (post-training).
- Pre-training is similar to Grok 3.
- Three variants:
- Pre-trained model with RL: 27% on Humanities Last Exam (similar to Gemini 1.5 Pro).
- With tool usage: Significant performance improvement on Humanities Last Exam.
- Multi-agentic system: Almost 50% on Humanities Last Exam.
Pricing and Availability
- "Grok Heavy" or "Super Grok Heavy" is expected to cost $300 per month.
- "Super Grok" is expected to cost $30 per month.
- Pricing is the most expensive LLM available.
- Context window: 256k tokens.
- Pricing is the same as Grok 3.
Future Developments
- A separate coding model is planned for release in a few weeks, focusing on low latency.
- Future plans include a multi-modal agent and a video generation model.
Independent Analysis
- Greg from ARC Foundation: Grok 4 is the top-performing publicly available model on ARC AGI 2.
- ARC AGI 2 requires models to learn skills and demonstrate them at test time.
- Mike Noob points out the unintuitive fact that models can perform well on Humanities Last Exam but poorly on ARC AGI 2.
- Artificial Analysis: Grok 4 achieves an intelligence index of 73, outperforming OpenAI, Anthropic, and Google models.
- This is the first time XAI has taken the lead in the intelligence index.
- The version of Grok 4 deployed on X (Twitter) may differ from the API version.
- Grok 4 is a reasoning model, but the XAI API does not share reasoning tokens.
- Grok 4 pricing is the same as Grok 3, more expensive than Gemini 1.5 Pro and Claude 3.
Key Benchmarks (Artificial Analysis)
- Grok 4 leads in the intelligence index and coding index (even before the specialized coding model is released).
- All-time high score in GPQA Diamond of 88%, surpassing Gemini 1.5 Pro (84%).
- 75 output tokens per second, slower than Claude 3, Gemini 1.5 Pro, but faster than Opus 4.
Notable Quotes
- "Somebody from XAI reached out to them that we want to test Gro 4 on AGI." - Greg, President of ARC Foundation
- "Perhaps the most unintuitive thing about the AI today is that an AI can simultaneously score 50 plus on humanity's last exam relatively hard for humans while only scoring 16% on RKGI2 which is relatively easy for humans." - Mike Noob
- "You can laugh at X folks sleeping in the office tents or grinding till 4:20 a.m. on weekends, but you have to admit they are the fastest moving AI lab out there."
Conclusion
Grok 4 represents a significant advancement in LLMs, achieving state-of-the-art performance on various benchmarks. The key to its success lies in increased RL compute and the use of tool usage and multi-agentic systems. While pricing is high, independent analyses confirm its leading position in the AI landscape. XAI's rapid progress highlights the importance of compute, data, and talent in training state-of-the-art models.
AI summaries can miss context or contain errors. Check important details against the original video.





