Key Concepts
AI Benchmarking, Frontier AI Intelligence, Reasoning Models, Open Weights Models, Cost Frontier, Speed Frontier, Output Tokens, Latency, Mixture of Experts (MoE), Model Sparsity, Inference Software Optimizations, Hardware Improvements, Compute Demand.
AI Progress & Multiple Frontiers
The AI landscape has experienced rapid advancements in the last two years, spearheaded by OpenAI's GPT models. While models like GPT-4, Gemini and Claude currently lead in overall intelligence as per the Artificial Analysis Intelligence Index (a composite index of seven evaluations), there exist other frontiers to consider beyond just peak intelligence. These frontiers involve trade-offs, meaning the most intelligent model isn't always the optimal choice. The presentation focuses on reasoning models, open weights models, cost, and speed.
Reasoning Models Frontier
Reasoning models, while offering greater intelligence, require significantly more output tokens compared to non-reasoning models. The quantity of output tokens directly impacts both latency and cost.
- Output Tokens: Reasoning models are significantly more verbose, using an order of magnitude more output tokens. For example, GPT-4 Turbo required 7 million tokens to run the intelligence index, while GPT-4o Mini High took 72 million, and Gemini 2.5 Pro took 130 million.
- Latency: API latency is also substantially higher for reasoning models. GPT-4 Turbo had a median latency of 4.7 seconds, whereas GPT-4o Mini High exceeded 40 seconds. High latency impacts user experience, especially in applications requiring responsiveness like chatbots. Studies have shown user drop-off correlating with application latency.
- Agentic Applications: The impact of latency is amplified in agentic applications involving multiple sequential queries. For example, 30 queries at 10 seconds each result in a 5-minute wait, whereas 30 queries at 1 second each take only 30 seconds. This difference dramatically influences application design.
- Example: A contact center application may tolerate 30-second response times, but 5-minute delays are unacceptable.
Open Weights Frontier
The gap in intelligence between open weights and proprietary models has significantly narrowed.
- Intelligence Gap: The difference in intelligence between open weights and proprietary models is less than ever before, with recent releases like Deepseek R1 being very competitive.
- Key Players: Leading open weights models originate primarily from China-based AI labs like Deepseek and Alibaba (with its Quen 3 series). Meta and Nvidia (with Neotron fine-tunes of Llama) are also significant contributors.
Cost Frontier
The cost of accessing AI intelligence varies significantly between models.
- Cost Differences: GPT-3 cost $2,000 to run the intelligence index. GPT-4 Turbo is approximately 30 times cheaper, and GPT-4 Turbo Nano is over 500 times cheaper than GPT-3.
- Cost Considerations: When building applications, consider the cost structure. An agentic application with 30 sequential API calls on a cheaper model might still be more cost-effective than a single query to GPT-3.
- Reasoning Token Cost: The cost per token isn't the only factor. The verbosity of models, particularly the reasoning tokens output during thinking, contributes to overall cost. Labs might not emphasize this, but these tokens are billed as output.
- Cost Trends: The cost of accessing GPT-4 level intelligence has fallen over 100 times since mid-2023 across all quality bands. New quality bands see rapid cost reductions within months of emergence.
- Future Planning: When building applications, consider what would be possible if cost were not a barrier, as cost reductions may make previously infeasible applications viable in the future.
Speed Frontier
The speed at which output tokens are received has increased dramatically since early 2023.
- Speed Increases: Models are grouped by intelligence level to account for the typical intelligence/speed trade-off. All intelligence levels have seen increased output token speeds.
- Example: Accessing GPT-4 level intelligence increased from around 40 output tokens per second in 2023 to over 300 tokens per second.
- Contributing Factors: Several factors contribute to increased speed:
- Model Sparsity: Mixture of Experts (MoE) models activate only a portion of parameters at inference time, reducing compute per token.
- Model Distillation: Distillation techniques, such as 8B distillations, improve efficiency.
- Inference Software Optimizations: Techniques like Flash Attention enhance performance.
- Hardware Improvements: Newer hardware like H100 and B200 offer significant speed gains. The B200 achieves over 1,000 output tokens per second.
House View on Compute Demand
Despite increased efficiency, reduced costs, and hardware improvements, demand for compute will continue to increase due to:
- Larger Models: Models like Deepseek have vast numbers of parameters (over 600 billion).
- Demand for Intelligence: The desire for greater AI intelligence remains insatiable.
- Reasoning Models: Verbose reasoning models require more compute at inference time.
- Agentic Applications: Sequential requests in agentic applications multiply compute demand.
Therefore, even with improvements in efficiency, overall compute demand is expected to rise.
Synthesis/Conclusion
The AI landscape presents multiple frontiers beyond just peak intelligence, each with trade-offs in output tokens, latency, cost, and speed. Reasoning models offer higher intelligence but at the cost of increased output tokens and latency. Open weights models are rapidly closing the intelligence gap with proprietary models. The cost of accessing AI intelligence is decreasing rapidly, and output speeds are increasing. Despite these improvements, compute demand is expected to continue to rise due to larger models, the demand for higher intelligence, and the increasing use of agentic applications. Understanding these trade-offs and considering the future landscape is crucial for effective AI application development.
AI summaries can miss context or contain errors. Check important details against the original video.