Gemini 3.0 (Riftrunner Fully Tested): The WORST Gemini 3 Checkpoint YET.

By AICodeKing

Share:

Key Concepts

  • Gemini 3 Checkpoints: Various experimental versions of Google's Gemini 3 model released on LM Arena.
  • LM Arena: A platform for testing and comparing different large language models.
  • RiftRunner: The latest Gemini 3 checkpoint discussed in the video.
  • X58 Checkpoint: A previously tested Gemini 3 checkpoint that the speaker highly prefers.
  • Quantization: A process of reducing the precision of model weights, often to reduce model size and improve inference speed, but potentially impacting performance.
  • Flash Model: A hypothesized variant of Gemini, potentially optimized for faster inference and live speech capabilities.
  • Parameter Count: A measure of a model's size, with larger numbers generally indicating more complex models.

RiftRunner: A New Gemini 3 Checkpoint on LM Arena

The video discusses the release of a new Gemini 3 checkpoint named "RiftRunner" on LM Arena. This follows previous checkpoints such as X58, 2HT, lithium flow, and ECPT. The speaker expresses fatigue with the continuous release of checkpoints rather than a full model release. RiftRunner is presented as a new release candidate for Gemini 3.

Performance Evaluation of RiftRunner

The speaker conducted tests on RiftRunner using various prompts on LM Arena's battle mode.

  • Floor Plan Generation: The output was described as "kind of fine" and "pretty bland," lacking the interactivity of the X58 checkpoint (e.g., furniture drifting, lighting). It was considered better than Sonnet's output but slightly worse.
  • SVG Panda Holding a Burger: The burger generation was "extremely good," but the panda was "not as great." Overall, it was considered one of the best all-rounded generations yet, though still surpassed by X58 and other checkpoints.
  • Pokeball in 3JS: The generation was "pretty good" and "super good," with an improvement of no sky background compared to previous generations.
  • Chessboard with Autoplay: This prompt resulted in a failure, marking the first instance of a Gemini 3 checkpoint model failing a question. The speaker noted this was "not a good look for the model."
  • Minecraft Clone in Kandinsky Style: The environment generation was "kind of fine," but the character's jumping behavior was flawed, causing it to "jump into oblivion" and roam in the sky.
  • Majestic Butterfly Flying in the Garden: This was highlighted as "one of the best generations yet," with "really very good-looking" animation and garden, and a "super amazing" butterfly.
  • CLI Tool in Rust: The generation was "good."
  • Blender Script: The script was "fine" but not as advanced as the X58 checkpoint, which included features like lighting and texture.
  • Math Questions: RiftRunner passed one math question but failed the other.
  • Riddle: The riddle was a "good pass," and unexpectedly, the model also generated an HTML page for the riddle, which the speaker found "very weird."

Comparative Analysis and Ranking

Based on these tests, RiftRunner was ranked in the fifth position among all tested checkpoints, scoring the lowest.

  • Improvement but Not a Leap: The speaker considers RiftRunner a "good improvement" but not a "3.5 Sonnet moment."
  • Performance Decline: It was noted that RiftRunner performed "a bit worse" than the ECPT checkpoint, which itself had already shown a performance decrease.
  • Comparison to Sonnet: RiftRunner scores approximately 15% above Sonnet, which is considered "great."
  • Comparison to X58: However, it scores about 14% lower than the "best X58 checkpoint."

Potential Reasons for Performance Differences

The speaker speculates on the reasons for RiftRunner's performance, considering:

  • Security Filters: The addition of security filters might impact performance.
  • Chat Use Case Tuning: Tuning the model more for chat use cases could lead to different output characteristics.
  • Quantization: This is a strong possibility, where the model's precision is reduced to optimize for size and speed. The speaker suggests it might be a "quantized variant."
  • Non-thinking or Low-thinking Variant: While considered, the speaker dismisses this as Gemini Pro models are generally auto-thinking and would likely have allocated thinking budgets.
  • Flash-based: The speaker also considers if it's flash-based, but finds this unlikely given the performance difference.

The speaker emphasizes the lack of official details about the model, making definitive conclusions difficult.

Future Expectations and Speculations

  • Official Launch: The speaker hopes for an official launch of the models rather than more checkpoints.
  • Access to X58: A desire is expressed for continued access to the X58 checkpoint, possibly in pro or ultra modes.
  • Ultra Model: The possibility of an "ultra model" is raised, potentially competing with models like Opus, but with concerns about cost.
  • Apple-Google Deal: A deal between Apple and Google is mentioned, involving a "1.2 trillion parameter model" expected to be Gemini 3. The speaker believes a "flash model" might be used for this due to its live speech capabilities and need for faster inference.
  • Parameter Estimates: If the flash model is 1.2 trillion parameters, the speaker estimates the Pro version could be around 2 trillion parameters.
  • Nano Banana Variant: The speaker also mentions anticipation for the "Nano Banana new variant," which shows promise in early tests.

Conclusion and Call to Action

The speaker reiterates their preference for the X58 checkpoint, feeling that the continuous release of checkpoints is "breaking my trust little by little." They urge Google to launch the models already. The video concludes with a call for viewer engagement through comments, subscriptions, Superthanks donations, and channel memberships.

Chat with this Video

AI-Powered

Load the transcript when you're ready to chat so the initial page stays lighter.

Ready to summarize another video?

Summarize YouTube Video