Insane 3D models, realtime AI video, new #1 open model, realtime AI worlds, Gemini 3 Flash: AI NEWS

By AI Search

Share:

AI Weekly Update: New Models & Advancements (Detailed Summary)

Key Concepts: Open-source AI models, 3D generation, video generation, image editing, real-time rendering, AI animation, multimodal AI, diffusion models, VAEs, agentic AI, video avatars, egocentric video generation.

I. Interactive 3D World Generation: Hunyen World 1.5

Hunyen World 1.5 is a new open-source AI capable of creating interactive 3D worlds in real-time. Users can navigate these environments using standard keys (WD or arrow keys). The generation isn’t pre-designed; the scene is created dynamically as the user moves. While quality isn’t perfect (some noise and artifacts are present), the real-time aspect is a significant achievement. The model supports prompt-based additions to the scene (e.g., adding smoking wood, darkening the sky, causing explosions). Hunyen World 1.5 outperforms competitors like Matrix Game and Gamecraft in reconstruction quality. It requires a CUDA GPU with at least 14GB of VRAM for local operation. The training code is planned for open-source release. The potential application for video game design is highlighted, envisioning AI-generated worlds responding to player input.

II. 3D Image Generation from 2D: Stereo Space

Stereo Space is an open-source AI model that generates 3D scenes from 2D images. It produces images viewable with red/green anaglyph glasses, creating a 3D effect. Alternatively, users can cross their eyes to perceive the 3D render. Stereo Space outperforms other 3D photo generators based on benchmark results. A free Hugging Face Space is available for testing. The model requires approximately 12GB of VRAM to run locally. The process involves breaking down the image into components for 3D viewing.

III. Long-Form Video Generation: Long V2

Long V2 addresses a limitation of many AI video generators – short video length. It can generate videos up to 5 minutes long while maintaining coherence and consistency. Benchmarks demonstrate Long V2’s superior performance in long video quality and consistency compared to competitors. The model utilizes a 14B base model and requires approximately 14GB of VRAM for local execution. The GitHub repository is publicly available.

IV. Real-Time Talking Avatar Generation: Real Video (ZAI)

ZAI, the creators of GLM, released Real Video, an AI capable of generating videos of a person talking in real-time. This is achieved by combining text-to-speech with avatar animation, including lip-sync and facial expressions. The system boasts a delay of approximately 2 seconds. It leverages the open-source video model 2.2 plus a self-forcing technique for accelerated generation. The model and instructions for local execution are available on GitHub.

V. Image Layering: Quen Image Layered (Alibaba)

Quen Image Layered, developed by Alibaba’s Quen team, separates a single image into editable layers (background, character, text, etc.), similar to Photoshop layers. This allows for selective editing of image components without affecting others. Users can modify individual layers (e.g., change background color, alter character appearance). The model requires significant resources (60GB total, 40GB for the Transformer model) and a high-end GPU, though community-driven compressed versions are anticipated. The model is open-sourced.

VI. Open-Source 3D Model Generation: Trellis 2 (Microsoft)

Trellis 2 is a new open-source 3D model generator, a significant upgrade from its previous version. It generates highly detailed 3D models from single 2D images, even inferring unseen sides of objects. It excels at rendering complex details like fur. Trellis 2 utilizes voxels (selective 3D pixels) and sparse compression VAE to efficiently represent 3D data, compressing it by a factor of 16 in each dimension. A free Hugging Face Space is available for online testing. Local execution requires an Nvidia GPU with at least 24GB of VRAM.

VII. Accelerated Video Generation: Turbo Diffusion

Turbo Diffusion significantly speeds up local AI video generation (100-200x faster) when used with models like 1 2.2 or 1 2.1. It can generate a 5-second video in as little as 2 seconds. It employs techniques like Sage Attention, Sparse Linear Attention (SLA), and RCM to achieve this speedup without substantial quality loss. The models are publicly available. Integration with Comfy UI is expected.

VIII. Character Animation with Movement Transfer: Scale

Scale is an open-source AI tool for animating characters by transferring movements from a reference video. It outperforms existing tools like One Animate in handling full-body movements and complex character shapes. It accurately captures movements, even with challenging actions. Scale utilizes 3D pose estimation for improved accuracy. The code is available on GitHub and integrated into Comfy UI via Swan video wrapper.

IX. Intrinsic Video Editing: VRBGX

VRBGX allows editing of intrinsic video properties: albido (color), normal (surface direction), irradiance (lighting), and material. Users can selectively modify these properties to alter the appearance of objects and the overall scene. The model is under development by Adobe, with plans to open-source the model weights, inference code, and training code.

X. Video Reshooting & Editing: Ray 3 Modify (Luma Labs)

Ray 3 Modify, from Luma Labs, enables reshooting existing videos in different styles or angles, adding visual effects, or replacing elements. It’s a powerful tool for AI-driven video editing, reducing the need for traditional techniques like mocap or green screens.

XI. Motion Control: Cling01

Cling01 released Motion Control, allowing users to control character movement using a reference video. It handles complex actions, lip-sync, and expressions. It supports reference videos up to 30 seconds long.

XII. SVG Text-to-Image: Cling SVG

Cling’s SVG text-to-image model generates images directly in visual space, bypassing the need for a VAE. While still experimental, it demonstrates promising results. The entire pipeline, including models and code, is open-sourced.

XIII. Video Generation: 1 2.6 (Alibaba)

Alibaba’s 1 2.6 generates videos with audio up to 1080p, with improved lip-sync and multi-shot capabilities. It includes a reference video feature. However, it’s considered a marginal upgrade and is closed-source.

XIV. Enhanced Video Generation: C-Dance 1.5 Pro (ByteDance)

ByteDance’s C-Dance 1.5 Pro excels in video generation quality and consistency, outperforming competitors in benchmarks. It’s available in the Streaming Up platform.

XV. Open-Source Reasoning Model: Mimo V2 Flash (Xiaomi)

Xiaomi’s Mimo V2 Flash is a powerful open-source model excelling in reasoning, agentic tasks, and coding. It performs competitively with top proprietary models like Claude 4.5 and GPT-5, despite being smaller in size (309 billion parameters). It’s the best open-source model for agentic coding, surpassing DeepSeek and Kimiku. The model is available on Hugging Face.

XVI. Multimodal AI: Gemini 3 Flash (Google)

Google’s Gemini 3 Flash is the most efficient model currently available, offering a balance of performance and cost. It excels in multimodal tasks (image, video, audio) and agentic coding. It boasts a 1 million token context window and is now the default model in the Gemini app.

Conclusion:

The AI landscape is rapidly evolving, with a surge of new open-source models and advancements in video and image generation. Key trends include real-time rendering, long-form video creation, enhanced character animation, and efficient multimodal AI. The open-sourcing of powerful models like Mimo V2 Flash and Trellis 2 is democratizing access to cutting-edge AI technology, while tools like Turbo Diffusion are making local execution more accessible. The focus is shifting towards more intuitive and versatile AI-powered video editing, promising a future where content creation is significantly streamlined.

Chat with this Video

AI-Powered

Load the transcript when you're ready to chat so the initial page stays lighter.

Ready to summarize another video?

Summarize YouTube Video