Why The Fix For AI's Spending Problem Is Not Good For OpenAI And Anthropic
By CNBC
Key Concepts
- Model Routing: The practice of directing AI tasks to different models based on complexity and cost—using high-end "frontier" models for difficult problems and cheaper, faster models for boilerplate or routine tasks.
- Tokconomics: The study and management of token consumption costs in AI, which have become a significant budgetary concern for enterprises.
- Agentic AI: AI systems (agents) capable of performing autonomous tasks, often triggering other agents or events without human intervention.
- Deskside Computing: A shift toward running smaller, efficient AI models locally on hardware (e.g., Mac Minis) at the user's desk to reduce latency and cloud costs.
- Jevons Paradox (in AI): The observation that as AI becomes more efficient and cheaper, overall consumption and demand increase, leading to higher total usage rather than cost savings.
- AI Productivity Guarantee: A vendor commitment (e.g., Cognition’s $10M guarantee) to ensure AI delivers measurable ROI, shifting the risk from the customer to the provider.
1. The Shift in Corporate AI Strategy
Corporate America is moving away from the "one-size-fits-all" approach of using the most expensive, powerful model for every task. This "old way" is compared to assigning a top-tier, highly paid engineer to perform simple tasks like resetting passwords. The "new way" is model routing, which optimizes for cost-efficiency.
- Cost Disparity: A complex task might cost $25 on a frontier model, while a routine task can be handled for under $1 by a smaller model.
- Industry Impact: Companies like Open Router have seen volume increase five-fold in six months, signaling that the era of picking a single model is ending.
2. ROI and Productivity Guarantees
As AI budgets are being exhausted in months rather than years, vendors are under pressure to prove value.
- Cognition’s Approach: CEO Scott Woo introduced an "AI productivity guarantee," where the company essentially puts its money behind the value its coding agent, Devon, provides.
- Measurement Methodology: Success is measured by "hours of engineering effort" saved and code successfully shipped to production, rather than vanity metrics like "lines of code" or "token counts."
- Case Study: Mercedes-Benz utilized Devon to complete migrations in 8 days that were originally forecasted to take 8 months.
3. Enterprise Perspectives: The Cisco View
G2 Patel, President and Chief Product Officer at Cisco, highlights that enterprises are currently navigating three phases: familiarity, proficiency, and efficiency.
- The Budget Reality: Cisco acknowledges that token usage is skyrocketing. They have had to reprioritize budgets—moving funds from traditional marketing (e.g., billboards) to AI infrastructure and token consumption.
- Network Strain: Agents are significantly more resource-intensive than humans, requiring 450% more network bandwidth to perform the same tasks.
- Infrastructure Super-Cycle: Cisco is seeing a 25% growth in campus/branch networking, driven by the need to support agentic workflows and the rise of deskside computing.
4. The Future of Model Usage and Security
- Local vs. Cloud: There is a growing trend toward running models locally on deskside hardware to bypass cloud costs and security concerns.
- Security Concerns: While some enterprises remain wary of foreign-hosted open-source models, the consensus is that security is managed through standard enterprise guardrails (QA, code review, and human-in-the-loop checkpoints) regardless of the model's origin.
- The "Neutral" Stance: Companies like Cognition emphasize their independence, arguing that being model-agnostic allows them to be "value-aligned" with customers by recommending the best model for the specific task, rather than pushing a single provider.
5. Notable Quotes
- Scott Woo (Cognition): "If you spend $5 on a task, that's fine as long as you're getting $20 of output out of it. It's not okay if you're spending $500 to get that $20 of output."
- G2 Patel (Cisco): "The biggest risk we have as an industry is if this gets too expensive where the cost of tokens is disproportionately higher than the value it generates... the cost per token going down is better for the model providers because when the cost goes down, people use it more."
Synthesis and Conclusion
The AI industry is entering a "maturation phase" where the focus is shifting from raw capability to economic sustainability. The "AI trade" is no longer just about the power of the model, but about the intelligent orchestration of multiple models. While frontier labs will continue to thrive due to the high value of complex, strategic tasks, the commoditization of "easy" tasks via model routing is inevitable. The long-term growth of the sector will be driven by the Jevons Paradox: as AI becomes cheaper and more efficient, its integration into every layer of the enterprise—from the data center to the desk—will expand, creating a massive, distributed ecosystem of coordinated intelligence.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

AI-Driven Multi-Document Correlation for Financial Compliance - Varsha Shah, Independent
AI Engineer

Z.AI And The Chinese Open Source Moment
CNBC

Việt Nam và bài toán xây dựng mô hình A.I tỷ đô: Có thực sự cần thiết? | Long Nguyễn
Spiderum

Stanford MS&E435 Economics of the AI Supercycle | Spring 2026 | Building AI Factories
Stanford Online

The AI Rollup Wave Transforming Main Street
CNBC

OpenAI on OpenAI: Stacie Faggioli, Business Finance Officer Applications, OpenAI
OpenAI

Stanford MS&E435 Economics of the AI Supercycle | Spring 2026 | Applications, Applied AI
Stanford Online