New top AI video models, Google dominates, top 3D model generator, Grok 4.1, GPT-Codex-Max: AI NEWS
By AI Search
Key Concepts
- Depth Anything 3: AI model for generating 3D maps from images/videos.
- Segment Anything Model 3 (SAM 3): AI for detecting, segmenting, and tracking objects in images/videos.
- SAM 3D: AI model for generating 3D models from images, with specialization in human bodies.
- Hunyuan Video 1.5: Open-source text-to-video and image-to-video generation model.
- Kandinsky 5: Open-source model for video and image generation/editing.
- Grok 4.1: XAI's language model.
- Gemini 3: Google's advanced multimodal AI model.
- Nano Banana Pro: Google's image generator and editor.
- GPT 5.1 Codeex Max: OpenAI's agentic coding model.
- Proactive Hearing Assistant: AI system for enhancing conversation clarity.
- Weather Next 2: Google DeepMind's advanced weather forecasting model.
- Anti-gravity: Agent-first IDE for AI-assisted coding.
- FizX Anything: AI for creating 3D models with estimated articulation from single images.
- Dr. Tulu: Open-source deep research agent.
- Part-aware Multimodal Model (PartXM LLM): AI for 3D model generation, editing, and grounding with part awareness.
- UniMo 2 Omni: Open-source omnimodal model understanding text, audio, images, and video.
New Open-Source Video Generators
This week saw the release of two state-of-the-art open-source video generators.
Hunyuan Video 1.5
- Description: Developed by Tencent, this is a lightweight model with 8.3 billion parameters, making it suitable for most consumer-grade GPUs.
- Capabilities:
- Generates high-quality videos from text prompts, capturing subtle emotions and anatomical correctness.
- Excels at rendering text within videos and simulating realistic physics.
- Follows specified camera movements (panning, pulling back, rising).
- Supports image-to-video generation, using an uploaded image as the starting frame for cinematic shots.
- Technical Details:
- Base model requires a minimum of 14GB VRAM with offloading.
- Outputs silent videos of 5-10 seconds duration.
- Natively outputs up to 720p resolution, with an enhancement for upscaling to 1080p.
- Performance: Claims to be better than the current leading open-source video generator, Wand 2.2, in instruction following, visual quality, structural stability, and motion effects, with a 17.12% higher win rate.
- Availability: Open-source with a GitHub repository containing instructions for local setup. Supported in Comfy UI, with potential for further speed-up using LightX2V.
Kandinsky 5
- Description: Another open-source model released this week, capable of generating both videos and images.
- Models Released:
- Video Pro: Most advanced, 19 billion parameters, generates videos up to 10 seconds.
- Video Light: Lightweight, 2 billion parameters, faster but with some quality sacrifice, runs on consumer GPUs.
- Image generation and editing models.
- Capabilities:
- Generates videos with good visual quality and motion, often physically correct.
- Can produce different perspectives, such as fisheye drone footage.
- Capable of cinematic shots and 3D animation styles.
- Limitations: Noticed artifacts and warping in high-action shots. Less attention and fewer community tools compared to Hunyan or Wand.
- Availability: Open-source with instructions on GitHub and a Comfy UI workflow available.
Advanced AI Models and Tools
Depth Anything 3
- Description: An AI model that creates a 3D map of a space from just a few images or a video.
- Capabilities:
- Captures detailed and coherent 3D scenes from various inputs, even chaotic, high-action videos with significant camera movement.
- Generates 3D animations, depth, and geometry of a scene.
- Can reconstruct 3D scenes from a large number of images.
- Performance: Faster, covers more space, and is more accurate than other similar models in terms of depth and geometry. Beats other state-of-the-art models in camera pose accuracy and reconstruction accuracy.
- Technical Details: The largest model, "DA3 nested giant large," has 1.4 billion parameters and is approximately 7GB in size, fitting on most consumer-grade GPUs.
- Availability: Open-source with a GitHub repository and interactive demos.
Segment Anything Model 3 (SAM 3)
- Description: An AI that detects, segments, and tracks objects in images and videos.
- Capabilities:
- Segments objects by clicking on frames or using text prompts (e.g., "zebra," "elephant").
- Can select multiple objects of the same category with a single prompt.
- Highly effective in segmenting and tracking specified objects in chaotic, high-action scenes.
- Performance: Blazing fast, detecting over 100 objects in 30 milliseconds on an H200 GPU. Outperforms competitor models in segmentation, concept segmentation, visual segmentation, object counting, and reasoning.
- Availability: Open-source with models released on Hugging Face (total size ~7GB) and a GitHub repository for local setup.
SAM 3D
- Description: A powerful 3D model generator that converts specified objects in an image into accurate 3D models.
- Capabilities:
- Generates 3D models of individual objects within a scene.
- Specializes in generating human bodies, including irregular poses and postures.
- Performance: Outperforms leading 3D generators like Trellis, Hunen 3D, Trippo, and High 3D Gen across benchmarks for 3D shape and texture preference, and accurately reconstructing human body meshes.
- Availability: Released with two GitHub repositories: one for general objects and one specifically for human bodies. An online demo is available.
Grok 4.1
- Description: XAI's latest top model, released this week.
- Performance: At the time of release, it was the best model on the LM Arena leaderboard for thinking and general performance, beating Gemini 2.5 Pro and GPT 5 Hive. It also led in emotional intelligence benchmarks and was strong in creative writing.
- Note: Described as an incremental improvement over previous versions.
Gemini 3
- Description: Google's new, highly advanced multimodal AI model.
- Performance: Dominates most leaderboards, ranking number one on LM Arena (text), WebDev, Vision, and Artificial Analysis.
- SweetBench Verified: Ranks number one for agentic coding, surpassing Anthropic's self-reported score.
- GeoBench: Ranks number one for guessing locations from images, even beating professional human players.
- Medical Images: Best in analyzing medical images, outperforming human radiology trainees.
- Simple QA: Ranks number one.
- Limitations: Has a relatively high hallucination rate.
- Availability: A full review video was previously released.
Nano Banana Pro
- Description: Described as the best image generator and editor available, released shortly after Gemini 3.
- Capabilities:
- Generates any character or famous person.
- Remasters old games.
- Analyzes medical images.
- Note: A full review video was previously released.
GPT 5.1 Codeex Max
- Description: OpenAI's best agentic coding model, released after Gemini 3.
- Focus: Specifically trained for agentic coding, designed for long-running, detailed tasks requiring multi-hour agent loops and autonomous execution.
- Performance: Scores significantly better than previous versions (5.1 Codeex High) on SuiLancer and Terminal Bench. Achieves 76.8% (High) and 77.9% (X High) on SuiBench Verified, outperforming Gemini 3 Pro (74.2%).
- Availability: Available on OpenAI's Codex platform for paid users.
Proactive Hearing Assistant
- Description: A useful AI system that helps people hear conversations more clearly by separating target voices from background noise.
- Capabilities: Automatically detects and isolates the voice of the person the user wants to hear, effectively removing other sounds.
- Use Cases: Beneficial for individuals with hearing difficulties in noisy environments.
- Availability: The dataset and training code have been released on GitHub, allowing users to install, download data, and train/evaluate the model.
Weather Next 2
- Description: Google DeepMind's most advanced and efficient AI weather forecasting model.
- Performance:
- Eight times faster than existing methods.
- Predicts weather with hourly resolution up to the hour.
- Generates hundreds of possible weather outcomes from a single starting point in less than a minute per prediction on a single TPU.
- Surpasses previous models on 99.9% of variables (temperature, wind, humidity).
- Takes hours on supercomputers using physics-based models for comparable predictions.
- Methodology: Uses a "functional generative network" (FGN) approach that injects noise for physical realism.
- Deployment: Being deployed across Google's tools, including Search, Gemini, Weather, and the Google Maps Weather API.
Anti-gravity
- Description: An agent-first IDE (Integrated Development Environment) or development platform.
- Capabilities:
- Allows users to command a team of AI agents that autonomously work on codebases for tasks like adding features, fixing bugs, or refactoring.
- Features a built-in browser for agents to load and test websites live, detect errors, and fix them autonomously. This is a key differentiator from other IDEs like Cursor.
- Availability: Currently in public preview and free to use. Available for Windows, Mac, and Linux. Powered primarily by Gemini 3, but allows switching between models.
- Note: Initial testing suggests it can be buggy, and free daily credits are consumed quickly.
FizX Anything
- Description: An AI that creates 3D models of objects from a single image, estimating articulation.
- Capabilities:
- Generates stationary 3D objects.
- Estimates object articulation (how it can move, e.g., opening/closing).
- Assets can be directly deployed into simulations for robot training.
- Methodology: Uses a new method that compresses tokens by 193 times, feeding them through a vision-language model to predict material, kinematics, description, and geometry.
- Dependencies: Based on Trellis (3D model generator) and Quen 2.5 (vision-language model).
- Availability: Code is released on GitHub with instructions for local setup. Requires a high-end GPU due to dependencies.
Dr. Tulu
- Description: An open-source deep research agent comparable to OpenAI's deep research models.
- Capabilities:
- Plans and executes tasks in multiple steps, calling tools like web search and data scraping.
- Gathers, compiles, and synthesizes information into a final answer.
- Performance: Despite being an 8 billion parameter model, it outperforms OpenAI's deep research and Perplexity Deep Research on Scholar QA v2 and Deep Research benchmarks. It closely matches closed proprietary models in evidence support and synthesis.
- Models: Two versions available: one trained on reinforcement learning (performs slightly better) and one on supervised fine-tuning.
- Availability: Released by the Allen Institute with a permissive license. GitHub repository provides setup instructions for local execution.
PartXM LLM (Part-aware Multimodal Model)
- Description: A multimodal model designed for 3D tasks, including generation, editing, and grounding, with part awareness.
- Capabilities:
- Generates 3D models as individual parts, enabling understanding and generation of complex structures.
- Allows for editing specific components of a model.
- Analyzes objects and reasons about them using natural language (e.g., counting breadsticks).
- Enables 3D object editing with natural language commands (e.g., replacing a basket liner).
- Availability: GitHub repository released, but code and models are not yet available.
UniMo 2 Omni
- Description: An open-source omnimodal model that understands and generates text, audio, images, and video.
- Architecture: Based on Quen 2.5 7B dense architecture, with a mixture of experts model trained on multimodal data.
- Capabilities:
- Understands and generates images, text, and speech.
- Can analyze and solve problems from images (e.g., bar graphs, math homework).
- Image generation quality is noted as similar to early Stable Diffusion or Flux.
- Can edit images (e.g., remove objects, remove rain).
- Features "ControlNet-like" capabilities for generating images from edge maps or depth maps.
- Understands both audio and video simultaneously (e.g., identifying dance moves from audio and video).
- Transcribes audio to text with impressive performance.
- Performance: On average, performs best among multimodal models of similar sizes.
- Technical Details:
- Unified encoding layer for various modalities into tokens.
- Omnimodality 3D rope layer for context tracking.
- Dynamic router for efficient data routing through experts.
- Decoder for outputting text, audio, or images.
- Availability: GitHub repository released with instructions for local setup. Different variants are available: Omni (handles all modalities, 87GB), Image (specialized for image generation/editing), and Text-to-Speech. The Omni model requires multiple GPUs.
Other Notable AI Releases
- Google Gemini 3: Described as a "freaking monster" that dominated leaderboards upon release.
- OpenAI's Best Coding Model: Quietly released after Gemini 3, named GPT 5.1 Codeex Max.
- GenSpark: A sponsor tool highlighted for its AI developer, designer, and slides agents, as well as AI inbox and teams features.
Synthesis and Conclusion
The past week has been exceptionally active in the AI landscape, marked by a surge of powerful new open-source models and significant advancements in existing AI capabilities. Key highlights include the release of two state-of-the-art open-source video generators, Hunyuan Video 1.5 and Kandinsky 5, offering impressive text-to-video and image-to-video functionalities. In the realm of 3D AI, Depth Anything 3 provides robust 3D mapping from images/videos, while SAM 3D excels at generating accurate 3D models, particularly of human bodies.
Google has made substantial contributions with Gemini 3, a multimodal model that has set new benchmarks across various tasks, and Nano Banana Pro, touted as the best image generator and editor. Weather Next 2 from Google DeepMind represents a significant leap in weather forecasting efficiency and accuracy. Furthermore, Google's Anti-gravity IDE aims to revolutionize coding with an agent-first approach.
OpenAI has also contributed with GPT 5.1 Codeex Max, their leading agentic coding model. The week also saw the release of practical AI tools like the Proactive Hearing Assistant, designed to improve auditory clarity, and FizX Anything, which creates articulated 3D models from single images.
The open-source community continues to thrive with Dr. Tulu, a highly capable deep research agent, and UniMo 2 Omni, a powerful omnimodal model that processes text, audio, images, and video. These releases underscore a trend towards more accessible, powerful, and specialized AI tools, pushing the boundaries of what's possible in generative AI, 3D modeling, coding, and data analysis.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

“I spent $50,000 self-hosting AI models. You should too.” - 0xSero
David Ondrej

Sakana Fugu (Fully Tested - V/S Fable): UHM... REALLY?
AICodeKing

Claude Sonnet 5, Mythos 6 ALREADY?, GPT-5.6 This Thursday, Sakana Fugu Beats Mythos, & More! AI NEWS
WorldofAI

Fable 5 Replacement Just Dropped: Fusion (Fable Level AI)
AI Revolution

OpenRouter Fusion: They OFFICIALLY CLAIM THAT THIS MODEL beats FABLE?
AICodeKing

Nex-N2 Pro IS GREAT! New Opensource Model Beats GPT 5.5, Opus 4,7, & Gemini 3.5? (Fully Tested)
WorldofAI

Claude Fable 5 IS INCREDIBLE! Greatest AI Model Ever! (Fully Tested)
WorldofAI