Codex-Spark: OpenAI Just Broke the Speed Limit (1,000 Tokens/s)

Prompt EngineeringAbout 5 min readFeb 14, 2026Watch original
THE SUMMARYAI-generated

Key Concepts

  • GPT-3.5 Codeex Spark: OpenAI’s new, ultra-fast model for real-time coding, powered by Cerebras hardware.
  • Gemini 3 Deep Think: Google’s model optimized for complex reasoning and challenging tasks.
  • Miniax M2.5 & GLM5: Recent open-weight model releases focused on coding and agentic engineering.
  • Agentic Users/Tasks: AI models capable of autonomous action and problem-solving, particularly in knowledge work.
  • Workhorse Models: Cost-effective, fast models designed for consistent performance in specific tasks.
  • Purpose-Built Hardware: Custom hardware designed to accelerate AI model inference, like Cerebras’ Scale Engine.
  • Context Window: The amount of text a model can consider at once (Spark: 128k tokens, GPT-3.5 Codeex: 200k tokens).
  • Round Trip Time: The time it takes for a request to be sent to a model and a response to be received.

OpenAI’s GPT-3.5 Codeex Spark & the Shift Towards Specialized AI Models

The AI landscape is rapidly evolving, moving away from a singular focus on Artificial General Intelligence (AGI) towards specialized models optimized for specific tasks, particularly coding. OpenAI’s recent release of GPT-3.5 Codeex Spark exemplifies this trend. This model, the first to run on custom, purpose-built hardware from Cerebras, achieves an impressive speed of 1,000 tokens per second – a significant leap forward for a frontier-level model. This speed is critical for real-time coding applications, like those found in codecs.

While Codeex Spark is a smaller version of the standard GPT-3.5 Codeex, resulting in some performance loss, OpenAI emphasizes that the “latest and greatest” model isn’t necessary for every task. The presenter argues that focusing on coding is strategically important because it enables tasks with “verifiable reward” and facilitates the development of “agentic users” – AI systems capable of autonomous action.

Currently, Codeex Spark is exclusively available to ChatGPT Pro users, with rate limits differing from the standard OpenAI API, due to hardware constraints with Cerebras wafer production. The model features a context window of 128,000 tokens, smaller than the 200,000 token window of the standard GPT-3.5 Codeex, but prioritizes speed over sheer intelligence. Demonstrations show Codeex Spark completing app development tasks significantly faster than the standard GPT-3.5 Codeex, which was still in the planning stages.

Contrasting Approaches: Gemini 3 Deep Think & Open-Weight Alternatives

Alongside Codeex Spark, other notable releases are shaping the AI landscape. Google’s Gemini 3 Deep Think is positioned as a model for tasks requiring “extreme level of reasoning and thinking.” It achieved a breakthrough on the RKGI2 benchmark, exceeding 80% accuracy – a previously challenging feat for large language models. However, Gemini 3 Deep Think is designed for accuracy over speed, making it more suitable for research-grade applications.

The presenter highlights a key contrast: OpenAI’s Codeex Spark prioritizes speed, while Gemini 3 Deep Think prioritizes intensive reasoning. This duality reflects a growing understanding that different tasks demand different model characteristics.

Furthermore, two recent open-weight model releases – Miniax M2.5 and GLM5 – are rapidly advancing the field. Miniax M2.5 currently leads benchmarks in agentic coding, surpassing GLM5 (released the day prior). While these open-weight models are capable, they are described as less “stable” and prone to inconsistent results, emphasizing the importance of individual experimentation.

The Rise of “Workhorse” Models & Purpose-Built Hardware

A central argument presented is the increasing importance of “workhorse” models – those offering a balance between intelligence, speed, and cost. The presenter suggests that models like Miniax M2.5 or GLM5 could even serve as viable substitutes for Codeex Spark as sub-agents, due to their impressive performance-to-cost ratio. This is particularly relevant for businesses where cost is a significant factor.

The success of Codeex Spark, powered by Cerebras’ Scale Engine 3, underscores the viability of “purpose-built hardware” for AI inference. This development poses a competitive challenge to Nvidia, which recently acquired Groc, indicating a broader industry trend towards specialized hardware solutions. OpenAI is targeting two key sectors with these releases: “long horizon reasoning and execution” and “real-time collaboration for rapid iteration.”

Technical Details & Performance Metrics

  • Tokens per Second: Codeex Spark achieves 1,000 tokens per second.
  • Context Window: Codeex Spark: 128,000 tokens; GPT-3.5 Codeex: 200,000 tokens.
  • Round Trip Time Reduction: OpenAI’s new responses API saw an 80% reduction in round trip time and a 50% reduction in time to first token with a persistent websocket connection.
  • RKGI2 Benchmark: Gemini 3 Deep Think was the first model to exceed 80% on this benchmark.
  • Hardware: Codeex Spark is powered by Cerebras’ Scale Engine 3.

Notable Quotes

  • “If you can get coding right, you can pretty much do any task that has verifiable reward.” – The presenter, emphasizing the strategic importance of coding in AI development.
  • “You don't need the latest and greatest model for every task.” – The presenter, advocating for a pragmatic approach to model selection.
  • “You want a compromise between speed and intelligence.” – The presenter, highlighting the need for balanced model characteristics.

Conclusion

The AI landscape is undergoing a significant shift towards specialized models optimized for specific tasks. OpenAI’s GPT-3.5 Codeex Spark, powered by Cerebras hardware, represents a major step forward in real-time coding capabilities. Alongside advancements from Google (Gemini 3 Deep Think) and the open-weight community (Miniax M2.5, GLM5), the industry is recognizing the importance of balancing speed, intelligence, and cost. The emergence of “workhorse” models and purpose-built hardware signals a move towards more practical and economically viable AI solutions, particularly for agentic tasks and knowledge work. The presenter concludes that 2024 is shaping up to be an incredibly dynamic year for AI innovation.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.