GPT-5.1 Codex (Fully Tested): This MODEL is ACTUALLY USEFUL! The best ALTERNATIVE to OPUS yet.

By AICodeKing

Share:

Key Concepts

  • GPT 5.1 Models: New lineup of OpenAI models, including "instant" (renamed chat model) and "thinking" (general model for API/complex tasks) variants.
  • Codeex Models: Upgraded models for coding tasks, with a "mini" and a larger, more capable version.
  • Instruction Following: A key improvement claimed for the GPT 5.1 models.
  • Benchmarks: Performance metrics used to evaluate models, with a note on their potential finickiness.
  • Pricing: Cost structure for the new models, with input and output token pricing.
  • Caching: Improvement in response API retention (24 hours) for cost-effectiveness in long-running tasks.
  • Kilo Code: A tool used for testing and interacting with the new models, requiring specific setup for optimal performance (JSON tool calling).
  • Agentic Tasks: Testing the models' ability to perform multi-step tasks or act as agents.
  • Vibe Coding: Creative coding where the model generates novel solutions, contrasted with strict prompt adherence.
  • Token Speed: A measure of model generation speed, with a significant difference noted between GPT 5.1 Codeex and other models like Sonnet.

OpenAI's New Model Lineup: GPT 5.1 and Codeex

OpenAI has released a new suite of models, including GPT 5.1 and upgraded Codeex models. The GPT 5.1 lineup features two versions:

  • GPT 5.1 Instant: This is essentially a rebranded version of their existing chat model within the GPT5 series.
  • GPT 5.1 Thinking: This is the general-purpose model intended for use via API or for handling complex computational tasks.

OpenAI claims a significant improvement in the instruction-following capabilities of these GPT 5.1 models.

Codeex Model Enhancements

The Codeex models have also received upgrades. There are now two versions:

  • Codeex Mini: Described as less capable.
  • Larger Codeex Model: Appears to be a robust and solid performer.

The transcript notes that benchmarks presented for Codeex sometimes use GPT 5.1 benchmarks, which the author finds acceptable, indicating a shift away from solely relying on the "SWE verified benchmark."

Pricing and Caching Improvements

The pricing structure for these new models remains consistent with previous offerings:

  • Larger Models (GPT 5.1 Thinking, larger Codeex): $1.50 per 1 million input tokens and $10 per 1 million output tokens.
  • Codeex Mini: Shares the same input token price but costs $6 per 1 million output tokens.

Caching has been improved for the responses API, now offering approximately 24 hours of retention. This enhancement makes it more cost-effective for long-running tasks.

Performance Testing and Observations

The author conducted personal tests using Kilo Code to evaluate the models' performance.

Initial Visual/Generative Tests:

  • Floor Plan: The generated floor plan was deemed "fine" with a sensible layout, but not extraordinary.
  • SVG Panda Eating a Burger: The result was "not great" and looked "pretty bad."
  • Pokeball in 3JS: This was rated as "insanely good," comparable to Gemini 3 level performance.
  • Chessboard: While a chessboard was generated, the autoplay functionality did not work, making it "not so good."
  • Minecraft Clone in Kandinsky Style: The output was described as "kind of fine." While it resembled a map rather than an actual playable game as requested, the visual quality was considered "really good."
  • Majestic Butterfly Flying in a Garden Simulation: This simulation was "kind of fine." The butterfly's wings were disproportionately large, but the motion and overall simulation were "pretty good," with a "nice" environment.

Coding and Reasoning Tests:

  • CLI Tool in Rust: Performed "fine."
  • Blender Script: "Just straight up doesn't work."
  • Math and Riddle Tests: Did not pass.

Benchmark Rankings and Comparisons:

  • GPT 5.1 Codeex: Scored highest among the tested variants, ranking 9th. It outperformed GLM4.6 but was surpassed by Claude. The author believes it performs better than GLM 4.6 in most tasks, though it can still be "finicky." This is considered a good improvement over the previous generation.
  • Codeex Mini: Performed "pretty bad," ranking 32nd.
  • General GPT 5.1 (High Reasoning Score): Ranked 16th.

Agentic Task Testing with Kilo Code

The author utilized Kilo Code for agentic task testing to maintain consistency with previous evaluations. To use the GPT 5.1 Codeex model within Kilo Code:

  1. Install Kilo Code.
  2. Navigate to Settings.
  3. Create a new profile.
  4. Select the GPT5 Codeex model.
  5. Crucially, in advanced settings, change the tool calling option to use JSON for improved error-free operation.

Results from Agentic Task Tests:

  • Movie Tracker App: The model completed the task, but the results were not impressive. The interface was functional but poorly organized, with everything on a single page, leading to a "not the best experience."
  • Godo Game: The task "didn't work" and produced numerous errors.
  • Goi Calculator: This was a success. The calculator worked "really good" and "the first time," with all keys functioning properly.
  • Open Code Repo Question: Did not work.
  • Spelt App: Worked but was "a bit buggy and not very usable."
  • Nux App: Did not work.
  • Rust App: Did not work.

The author notes that these agentic task failures contribute to the GPT 5.1 Codeex's 9th position ranking, slightly above GPT5 Codeex (which seems to be a typo and likely refers to a different model or a previous iteration). The author agrees that it is "slightly better for sure" than previous iterations.

Overall Assessment of GPT 5.1 Codeex

The author considers the GPT 5.1 Codeex model to be "a good model by OpenAI."

  • Strengths:
    • Excellent for planning tasks, showing improvement in this area.
    • Beneficial for integration into existing codebases.
  • Weaknesses:
    • Not ideal for "vibe coding" (creative, emergent coding) as it adheres too strictly to the prompt, lacking independent creative input.
    • Performance is slow, operating at approximately 18 tokens per second, which is significantly slower than models like Sonnet (around 80 tokens per second). This slowness makes it "a bit unusable" for tasks requiring extended generation time, such as a 30-minute completion for a pair programming task.

Due to its speed limitations, the author primarily uses this model for planning and debugging, rather than as a real-time pair programmer.

Conclusion and Future Outlook

The GPT 5.1 and Codeex models represent an advancement from OpenAI, particularly in instruction following and coding capabilities. While the larger Codeex model shows promise, especially for planning and integration, its slow generation speed is a significant drawback for interactive coding. The author expresses hope for future speed improvements. The author invites user feedback and encourages subscriptions and donations.

Chat with this Video

AI-Powered

Load the transcript when you're ready to chat so the initial page stays lighter.

Ready to summarize another video?

Summarize YouTube Video