Stanford MS&E435 Economics of the AI Supercycle | Spring 2026 | Infrastructure, Capstone Case

Stanford OnlineAbout 4 min readMay 29, 2026Watch original
THE SUMMARYAI-generated

Key Concepts

  • Industrial Compute: The end-to-end management of AI infrastructure, including chips, memory, networking, power, cooling, and data center construction.
  • Agentic Workloads: AI systems that move beyond simple "one-shot" inference to "closing the loop"—reasoning, using tools, iterating, and taking actions to complete complex tasks.
  • Scaling Laws: The observation that increasing compute, data, and model size leads to predictable improvements in intelligence; now expanded to include pre-training, post-training (RL), and synthetic data generation.
  • Compute Graph: The complex, directed acyclic graph (DAG) of inference calls, tool usage, and environment interactions required for modern AI agents.
  • Heterogeneous Compute: The shift from relying solely on GPUs to using a mix of specialized accelerators (e.g., Cerebras for fast inference) and CPUs to optimize for specific tasks.
  • Prefill Latency: The time required for a model to process the entire context (e.g., a large codebase) before generating the first output token.

1. The Economics of the AI Super Cycle

Professor Kati (OpenAI) emphasizes that revenue for frontier labs is a lagging indicator of compute capacity. OpenAI’s strategy is to maximize compute availability to ensure researchers remain unconstrained.

  • Correlation: OpenAI has historically tripled compute capacity year-over-year, which has directly correlated with revenue growth.
  • Inference Dominance: While training was the initial focus, inference now represents the "super majority" (80%+) of compute usage, driven by RL, synthetic data generation, and product usage (e.g., ChatGPT, Codex).
  • Monetization Strategy: OpenAI aims to lower the cost per token, increase token intelligence, and improve model efficiency (requiring fewer tokens to complete a task) to maximize accessibility.

2. Infrastructure Challenges and "Stargate"

Building at the gigawatt scale is an operational challenge that extends far beyond signing chip contracts.

  • The Supply Chain: Sourcing involves a massive, interconnected chain: power generation, cooling, land, networking, and specialized chips.
  • Grid Impact: A gigawatt-scale data center can cause massive energy fluctuations. Infrastructure must be redesigned to prevent grid instability.
  • Concentration vs. Distribution: Due to the high cost of building smaller clusters and the need for massive human capital, OpenAI favors large, concentrated compute clusters over distributed edge computing.
  • The "Stargate" Philosophy: The focus is shifting from "amount of compute" to "time to compute." Speed of deployment is the primary operational metric.

3. The Evolution of Agentic Workloads

The transition from chatbots to agents has fundamentally changed the compute requirements:

  • From Chat to Agency: Chatbots are passive; agents are active. Agents "close the loop" by spinning up VMs, searching databases, and iterating on outputs.
  • Latency Bottlenecks: The "prefill" phase (paging in large contexts like 400k tokens) creates a 400–500ms latency floor. Improving hardware (e.g., using Cerebras) exposes inefficiencies in the rest of the software stack, necessitating constant engineering to shave off milliseconds.
  • Human-in-the-Loop: The goal is to make the AI so fast that the human becomes the bottleneck, keeping the user in a state of "flow."

4. Market Perspectives and Future Outlook

  • The "Layer Cake" Value: Currently, value is concentrated in the infrastructure layer (chips, energy, data centers). Over time, value will likely migrate to the application and platform layers, similar to the mobile revolution.
  • The Recursion of AI: AI is increasingly being used to design the next generation of chips and low-level software, which is essential to shortening the current 3-year chip design cycle.
  • The Choke Point: The most significant structural bottleneck is the limited fab capacity at the TSMC/Samsung/Intel level, and even deeper, the reliance on ASML lithography machines.
  • Advice for Students: Kati encourages students to go "long" on the lowest layers of the stack (materials, power, transistors, and foundational infrastructure), noting that these areas are currently underserved and critical for long-term sustainability.

5. Notable Quotes

  • "Revenue is basically a lagging indicator for frontier lab companies." — Professor Kati
  • "It’s no good if you build the intelligence but you can’t really deliver it at scale." — Professor Kati
  • "Today, apps are a crutch to get to an outcome." — Professor Kati (on why he is bearish on simple model wrappers).
  • "The world needs a much more resilient compute supply chain. It is dangerous for the world to be single-threaded on any one component." — Professor Kati

Synthesis

The AI super cycle is shifting from a "training-heavy" phase to an "inference-heavy" agentic phase. This transition requires a move toward heterogeneous compute architectures and a massive expansion of foundational infrastructure (energy, fabs, and cooling). While the current market is dominated by GPU-centric infrastructure, the future will be defined by specialized accelerators and AI-driven system design. The ultimate goal is to make intelligence so accessible and efficient that the human user—not the compute—becomes the primary constraint in the workflow.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.