AI Demand Is Inflated And Only Anthropic Is Being Realistic
By CNBC
Key Concepts
- Tokens: The fundamental unit of AI consumption, representing the text or code processed by a model.
- AI Agents: Autonomous systems capable of browsing the web, writing code, and executing multi-step tasks over extended periods.
- Token Maxing: A counterproductive corporate practice of incentivizing employees to maximize AI usage volume rather than output quality.
- Inference Budget: The financial allocation for running AI models (inference) versus training them.
- Cone of Uncertainty: The gap between current massive infrastructure investment and the actual, unverified future demand for AI services.
- Per-Token Billing: A pricing model where costs are tied directly to usage volume, replacing flat-rate subscriptions.
The AI Demand Signal Crisis
The current AI investment cycle is characterized by massive capital expenditure on infrastructure, yet there is a growing disconnect between this spending and actual, sustainable demand. While major tech companies are pouring billions into data centers and chips, many lack a clear understanding of the financial risks associated with AI consumption.
1. The Shift from Chatbots to Agents
The transition from simple chat interfaces to autonomous AI agents has fundamentally changed the cost structure of AI.
- Consumption Disparity: A simple chat query costs a few hundred tokens. In contrast, AI agents can run in the background for hours, consuming millions of tokens.
- Budgetary Impact: Companies are struggling to forecast these costs. For example, the Uber CTO reported that AI coding tools exhausted the company’s full-year AI budget by April. Goldman Sachs Research indicates that companies are overrunning initial inference budgets by orders of magnitude.
2. The "Token Maxing" Problem
Corporate culture is currently misaligned with economic reality. Companies like Meta and Shopify have implemented leaderboards to track AI adoption, but these metrics often measure volume rather than productivity.
- Gaming the System: As noted by the CEO of Databricks, tracking usage instead of output encourages employees to "burn money" to climb leaderboards.
- Jensen Huang’s Perspective: The Nvidia CEO famously stated he would be "deeply alarmed" if a $500,000 engineer didn't consume at least $250,000 in tokens, highlighting a philosophy that equates high consumption with high productivity—a metric that experts argue is easily gamed.
3. The Unsustainability of Flat-Rate Pricing
The "all-you-can-eat" subscription model (e.g., ChatGPT Pro) is becoming economically unviable as agent-based workloads increase.
- The Cost Gap: Estimates suggest a $200 flat-rate plan could result in $2,000 to $5,000 in actual compute costs.
- Anthropic’s Pivot: Anthropic is leading a shift away from flat-rate plans toward per-token billing. By cutting off third-party tools that relied on unlimited subscriptions, they are forcing enterprise customers to pay for actual usage, thereby creating a more accurate and sustainable demand signal.
4. Infrastructure and the "Cone of Uncertainty"
The entire AI investment thesis—justifying hundreds of billions in chip sales and 30-gigawatt data center capacities—relies on the assumption of exponential growth in usage.
- The Risk of Overinvestment: Data centers require 1–2 years to build. Companies are making billion-dollar bets on demand that has not yet materialized.
- The "Ruinous" Gap: If the demand is inflated by "token maxing" or unsustainable flat-rate models, the infrastructure being built may be sized for a market that does not exist. As noted in the transcript, if a company is off by even a couple of years in their demand projections, the financial consequences could be ruinous.
Synthesis and Conclusion
The AI industry is currently experiencing a "broken" demand signal. While companies like Nvidia and major data center providers are betting on massive, sustained growth, the underlying usage patterns are often driven by inefficient corporate policies and unsustainable pricing models.
Anthropic stands out as the only major lab currently pricing for "real" demand. By moving to per-token billing and prioritizing verifiable ROI over vanity metrics, they are positioning themselves for a more stable future. As the market matures and companies move toward IPOs, those that can demonstrate clean, per-token data and a clear understanding of their return on investment will likely outperform those that relied on the "cool factor" and unchecked infrastructure spending. The core takeaway is that efficiency—using the right model for the right task—must replace the current culture of indiscriminate consumption.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

Stanford CS153 Frontier Systems | Building the Frontier Ecosystem
Stanford Online

'Things are going to be okay, in Canada and the U.S.': Thorne
BNN Bloomberg

I'M OUT: The $11 Trillion AI Bubble is Breaking!
Steven Van Metre

South Korea bets big on AI with nearly a trillion dollars of investment • FRANCE 24 English
FRANCE 24 English

The Bubble is Bursting... (Emergency Update)
Bravos Research

The AI Bubble Just Ended - Without Popping
Heresy Financial

AI Market Volatility, Europe Heat Wave, Venezuela Quakes Damage | Bloomberg This Weekend: June 27
Bloomberg Television