NEW GPT-4.1: POWERFUL Coding LLM! Beats Claude 3.7 and Gemini 2.5 Pro (Fully Tested)

WorldofAIAbout 3 min readApr 15, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

GPT-4.1, GPT-4.1 Mini, GPT-4.1 Nano, long context window (1 million tokens), coding performance, instruction following, latency, pricing, Swaybench verify test, RAG (Retrieval-Augmented Generation), Gemini 2.5 Pro, function calling, rate limits, 3.js, SVG generation.

GPT-4.1 Model Series Overview

OpenAI has launched the GPT-4.1 model series, available only through an API, consisting of three models:

  • GPT-4.1: The flagship model, excelling in coding, instruction following, and long context performance.
  • GPT-4.1 Mini: A lighter model with lower latency and cheaper pricing.
  • GPT-4.1 Nano: The fastest and most efficient model, ideal for autocompletes, classification, and large document processing.

These models outperform GPT-4 Omni and GPT-4 Omni Mini across various benchmarks. A key feature is the support for up to 1 million tokens of context, effectively addressing the "lost in the middle" issue.

Performance Benchmarks and Improvements

  • Coding: GPT-4.1 achieves a 54.6% score on the Swaybench verify test, a 22% improvement over GPT-4 Omni.
  • Instruction Following: Demonstrates state-of-the-art performance.
  • Long Context: Handles large contexts effectively.
  • GPT-4.1 Mini: Beats GPT-4 Omni in several benchmarks with nearly 50% lower latency and 83% cheaper pricing.

Pricing Details

The pricing for the GPT-4.1 series is as follows (per 1 million tokens):

  • GPT-4.1: Input - $2, Output - $8
  • GPT-4.1 Mini: Input - $0.40, Output - $1.80
  • GPT-4.1 Nano: Input - $0.10, Output - $0.40

Strengths and Use Cases

GPT-4.1 is an all-rounded coding model, proficient in both front-end and back-end tasks. It is faster and cheaper than GPT-4 Omni, with 80 times the speed. It is also a strong vision model, making it suitable for general use cases beyond coding.

The model's strengths include:

  • Long Context Window: Ideal for large code bases and documents, potentially eliminating the need for RAG in many cases.
  • Faster Responses: Provides quicker outputs compared to Gemini 2.5 Pro.
  • No Rate Limits: Avoids throttling issues experienced with other models.
  • Strong Function Calling: Reliable tool use.

It is positioned as a better choice than Cloud 3.5 Sonnet across every benchmark and is also cheaper.

Comparison with Gemini 2.5 Pro

While GPT-4.1 is a solid upgrade, the presenter believes it is not cheaper or better than Gemini 2.5 Pro in terms of reasoning capabilities. However, GPT-4.1 excels in scenarios where speed, long context, and zero throttling are critical.

Coding Task Tests

The video includes tests comparing GPT-4.1 and Gemini 2.5 Pro on various coding tasks:

  1. Frontend Generation (Monthly Income and Expenses Tracker): Both models generated responsive frontends, but neither was fully functional. Both models received a pass.
  2. TV Screen Simulation: Both models successfully generated code for a TV screen simulation with multiple channels. The presenter preferred the animation generated by GPT-4.1. Both models received a pass.
  3. SVG Butterfly Generation: Both models generated SVG representations of a butterfly. The presenter preferred the Gemini 2.5 Pro output due to better symmetry in the antenna and body. Both models received a pass.
  4. Tetris Game in 3.js: GPT-4.1 generated a functional Tetris game with the correct interface, while Gemini 2.5 Pro's output had issues. Both models received a pass.

Conclusion

GPT-4.1 is a solid, lightweight upgrade from OpenAI, offering a larger context window, faster responses, and strong function calling. It is particularly useful for processing larger documents and excels in scenarios where speed and long context are paramount. While not necessarily superior to Gemini 2.5 Pro in all aspects, it presents a compelling option for developers, especially when dealing with large codebases or documents. The presenter is looking forward to seeing it being added to chat GBT.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.