Infinite 3D worlds, long AI videos, realtime images, game agents, character swap, RIP Udio - AI NEWS
By AI Search
Key Concepts
- Long Cat Video: Open-source video generation model by Muan, capable of text-to-video, image-to-video, and video continuation for extended lengths.
- Emu 3.5: Open-source multimodal AI model from Google, excelling in language understanding, image generation, and editing.
- World Grow: AI for generating infinite, realistic 3D worlds with coherent geometry using a "building block" approach.
- Kimmy Linear: Hybrid linear attention model by Moonshot AI, designed for efficiency and performance with long context sequences.
- ChronoEdit: Nvidia's AI image editor that uses a video-based approach to apply edits over time.
- Sora 2 Character Cameos: Feature allowing insertion of fictional characters into videos.
- Artvar: OpenAI's agentic security researcher for finding, verifying, and fixing software vulnerabilities.
- Emergent: Autonomous AI agent for creating research reports, full-stack apps, and landing pages via prompts.
- Nitro E: Real-time image generator for AMD GPUs, prioritizing speed over quality.
- Pomelli: Google's AI tool for generating marketing creatives, potentially replacing design agencies.
- Mocha: AI video tool for character swapping in existing videos, maintaining motion and expressions.
- SVG (Self-Supervised Representation for Visual Generation): Quaiou's novel approach to image/video generation that removes the VAE step for improved speed and understanding.
- Neo1X Robot: Humanoid robot by 1X Technologies for home use, currently relying on teleoperation.
- Unitere G1 Robot: Humanoid robot demonstrating impressive strength by pulling a 1,400 kg car.
- Cuavo 5 Robot: Leju Robotics' modular humanoid robot with interchangeable limbs and Huawei's Pangu model integration.
- IGGT (Instance Grounded Geometry Transfer): AI for reconstructing 3D meshes and understanding object shapes from multiple images.
- Game Tars: Bite Dance's AI agent capable of autonomously playing any video game with human-like inputs.
- Greg (Group Relative Attention Guidance): AI tool for editing images with text prompts, allowing control over edit intensity and preserving scene integrity.
- Udio: AI music generator facing download restrictions due to a partnership and licensing agreements with Universal Music Group.
- Minimax Music 2.0: AI music generator for creating songs from text prompts, with optional lyrics.
- Highaw 2.3: Minimax's AI video model known for its physics and world understanding.
- Minimax M2: Open-source AI model claimed to be the number one open-source model, comparable to leading closed models.
- Foley Control: Stability AI's tool for generating audio for silent videos.
Long Cat Video: Extended Video Generation
Muan, a Chinese food delivery company, has released Long Cat Video, a highly flexible open-source video generation model. It supports text-to-video, image-to-video, and video continuation, enabling the creation of videos exceeding several minutes in length. The model demonstrates strong physics understanding, generating anatomically correct figures and accurate skateboard tricks. It can produce videos in various styles, including Disney Pixar and anime. For image-to-video, it can animate static images, useful for product commercials. The video continuation feature allows extending existing videos seamlessly, maintaining consistency of objects and characters over extended durations, a common flaw in previous models.
Technical Details:
- Model Size: 13.6 billion parameters.
- Output Resolution: 720p.
- Frame Rate: 30 frames per second.
- Availability: Models released on HuggingFace, with instructions on GitHub for local setup.
Emu 3.5: Multimodal AI for Image and Text
Emu 3.5 is an open-source multimodal model from Google that combines language and world understanding with advanced image analysis, generation, and editing capabilities. It can provide step-by-step instructions for tasks like sculpting or cooking, accompanied by coherent, generated images for each step. Emu 3.5 can also take an image as input and continue its generation, useful for creating storybooks with illustrations or figurines.
Key Features:
- Image Editing: Capable of tasks like clothes swapping and generating multiple references from input photos.
- World Understanding: Can solve visual puzzles and change photo perspectives.
- Object Removal: Can remove handwriting, watermarks, and obstructions.
- Performance: Claims to outperform leading image models like Quinn Imageedit, GPT, and Nano Banana on various benchmarks.
- Availability: Models released, with setup instructions on GitHub.
World Grow: Infinite 3D World Generation
World Grow is an AI that generates realistic and geometrically coherent infinite 3D worlds. Unlike traditional methods like Gaussian splatting, which can result in missing areas, World Grow uses a "building block" approach, akin to Lego, to construct scenes. This ensures structural and geometric consistency.
Methodology:
- Scene Block Collection & Preparation: Gathers and prepares building blocks for rooms, outdoor areas, and furniture layouts, converting them into structured 3D data.
- 3D Block Inpainting: Fills in any gaps, ensuring walls, floors, and lighting align correctly.
- Fine Structure Refinement: Adds realistic textures, materials, and lighting for enhanced detail and quality.
The block-based generation allows worlds to expand consistently, creating an explorable, ever-growing 3D environment. The code is planned for public release, with pre-trained weights and inference pipelines anticipated.
Kimmy Linear: Efficient Long Context Transformer
Kimmy Linear, developed by Moonshot AI (creators of Kimmy K2), is a hybrid linear attention model designed to improve the efficiency and performance of transformer models, especially for very long context sequences. It addresses the exponential increase in computation and memory required by traditional multi-head attention for long prompts.
Key Innovation:
- Kimmy Delta Attention: A refined version of gated delta that replaces the traditional multi-head attention step. It incorporates a linear attention mechanism, significantly reducing computation and memory usage.
- Efficiency Gains: Claims to reduce memory requirements by up to 75% and be six times faster in decoding compared to full attention models.
- Context Length: Supports a context length of 1 million tokens, far exceeding most current top AI models.
- Parameters: 48 billion parameters with only 3 billion active, making it potentially runnable on high-end consumer GPUs.
- Availability: Open-sourced on HuggingFace, with setup instructions on GitHub.
ChronoEdit: Video-Based Image Editing
ChronoEdit by Nvidia is an AI image editor that allows text-prompted edits. Its unique approach is built upon a video model, simulating the editing process as a short video. This allows for a more nuanced application of edits, visualizing the transformation over time.
Process:
- Instead of directly generating a final image, ChronoEdit generates a short video demonstrating the edit's progression.
- Examples include transforming an image into a figurine, rotating a character's pose, or changing food items.
- The model is available for local use via GitHub and has a free Hugging Face space for online testing.
- Requirement: At least 34 GB of VRAM.
Sora 2 Character Cameos
OpenAI's Sora 2 has introduced a Character Cameos feature. This allows users to insert fictional characters (pets, objects, non-realistic characters) into their videos by uploading a short video of the character. For realistic humans, the existing selfie video verification process is still required to ensure consent for deepfake generation.
Artvar: AI Security Researcher
Artvar by OpenAI is an agentic security researcher designed to find, verify, and help fix software vulnerabilities.
Workflow:
- Code Repository Analysis: Analyzes the codebase for potential issues.
- Vulnerability Scanning: Identifies vulnerabilities.
- Verification: Reproduces vulnerabilities in a sandbox environment.
- Fix Generation: Uses OpenAI Codex to suggest fixes.
- Human Review: The proposed fix is submitted to a human for review.
- Pull Request: Approved fixes are submitted as pull requests.
Effectiveness: OpenAI claims Artvar has identified 92% of known and synthetically introduced vulnerabilities during its internal use. Access is currently through a private beta for select partners.
Emergent: Autonomous AI Agent for Creation
Emergent is an autonomous AI agent that allows users to create various outputs, including research reports, full-stack applications, and landing pages, simply by providing a prompt.
Examples:
- Building a full-stack habit tracking app with a user-friendly interface.
- Generating a comprehensive financial analysis report for a company like Nvidia.
- Creating a responsive and visually appealing landing page for a marketing agency.
Emergent handles both front-end and back-end development automatically, making it accessible to non-technical users.
Nitro E: Real-time Image Generator for AMD GPUs
Nitro E is a new, efficient, and lightweight real-time image generator designed for consumer AMD GPUs.
Key Features:
- Speed: Generates six images per second on a consumer GPU.
- Model Size: Only 304 million parameters, making it very efficient.
- Training: Takes approximately 1.5 days to train on eight AMD GPUs.
- Performance: Can output nearly 19 images per second on a single AMD Instinct GPU.
- Availability: Open-sourced on HuggingFace (under 4 GB), with setup instructions on GitHub.
- Trade-off: Prioritizes speed and efficiency over image quality, with showcased images appearing less detailed than older models like Stable Diffusion 1.5.
Pomelli: Google's AI for Marketing Creatives
Pomelli by Google is a free AI tool designed to assist in creating marketing creatives. It has the potential to significantly impact marketing and design agencies.
Workflow:
- Website Scan: Users input their website, and Pomelli extracts branding elements like fonts, colors, and images.
- Campaign Idea Generation: Based on extracted branding, it generates campaign ideas.
- Creative Production: Automatically creates product photos, social media posts, and other marketing materials.
- Customization: Users can prompt for specific campaigns (e.g., "Facebook campaign for Black Friday up to 50% off") and edit generated creatives, including colors, leveraging extracted branding.
Pomelli aims to automate design processes, potentially competing with tools like Canva.
Mocha: Advanced Character Swapping in Video
Mocha is a powerful AI video tool that allows users to swap out characters in existing videos while preserving the original motions, gestures, and facial expressions.
Key Advantages:
- High Fidelity: Claims to offer better quality and higher fidelity than similar tools like Wan Animate.
- Background and Lighting Preservation: Excels at maintaining the original background and lighting conditions, as well as white balance.
- Character Appearance Transfer: Accurately transfers the appearance of the new character.
- Subtitle Handling: Can swap characters even with subtitles overlaid on the video.
- Availability: Models are available on HuggingFace, with setup instructions on GitHub. A ComfyUI workflow has also been developed.
SVG: Reinventing Visual Generation
SVG (Self-Supervised Representation for Visual Generation) by Quaiou (creators of Clling) is a new approach to image and video generation that eliminates the need for a Variational Autoencoder (VAE). VAEs are typically used to compress inputs into a latent space for diffusion models, but they can suffer from "semantic entanglement," making it difficult to distinguish between objects.
SVG Methodology:
- Removes VAE: Eliminates the VAE component.
- Integrates Dino: Uses Dino, an existing visual model, for object segmentation and understanding.
- Residual Encoder: Adds a small encoder for fine details.
Benefits:
- Improved Understanding: Better at world understanding and distinguishing objects.
- Speed: Up to 62 times faster in training and 35 times faster in inference compared to VAE-based models.
- High Fidelity: Generates high-quality images.
- Availability: Code released on GitHub for local experimentation.
Humanoid Robot Updates
Neo1X Robot (1X Technologies)
The Neo1X is a humanoid robot designed for home use, aiming to automate chores and provide assistance. It is priced at $20,000 with a $500/month subscription and is available for pre-order. However, demonstrations heavily rely on remote teleoperation by humans using VR headsets and controllers. Reports indicate limitations such as slow task completion and instability, with little evidence of true autonomy. The announcement is largely considered hype due to the current reliance on human control.
Unitere G1 Robot
The Unitere G1 robot demonstrates impressive strength by pulling a 1,400 kg car. This feat is achieved through the Thor algorithm (Towards Human-Level Whole Body Reactions), which enables the robot to adjust its posture for maximum pulling efficiency. While the car's wheels and neutral gear reduce the actual force required, the robot's ability to coordinate its entire body and adjust movements is highly impressive.
Cuavo 5 Robot (Leju Robotics)
The Cuavo 5 by Leju Robotics is a modular humanoid robot with swappable limbs (including feet or wheels for rolling). It boasts over 8 hours of battery life and a 20 kg load capacity. Originally designed for industrial settings, it is now being adapted for home use. It integrates Huawei's Pangu model as its "brain" for perception, decision-making, and control.
IGGT: 3D Reconstruction and Understanding
IGGT (Instance Grounded Geometry Transfer) is an AI that reconstructs 3D meshes and understands object shapes from multiple images taken at various angles. It can also identify and segment objects within the scene.
Key Features:
- Unified Model: Predicts and generates reconstruction, understanding, and tracking simultaneously within a single transformer model.
- Semantic Understanding: Unlike previous methods that only generated 3D scenes, IGGT understands the objects and shapes within them.
- Performance: Outperforms other 3D reconstruction models across various metrics.
- Availability: Codebase and pre-trained models are planned for release.
Game Tars: Autonomous Video Game Agent
Game Tars by Bite Dance is an AI agent designed to play video games using human-like inputs (visuals, keyboard, mouse). It can autonomously play any video game without prior knowledge of its rules.
Capabilities:
- Real-time Play: Lightweight and efficient enough to process gameplay and execute controls in real-time.
- Generalist Agent: Learns and adapts to any game, even those not seen during training.
- Performance: Outperforms models like GPT-5 and other top models in open-world tasks in Minecraft, FPS games, and web/simulator games.
- Potential Application: The technology could be applied to robotics for real-world tasks that robots have not encountered before, enhancing their capability and autonomy.
- Documentation: A technical paper is available for detailed information.
Greg: Gradient-Based Image Editing
Greg (Group Relative Attention Guidance) is an AI tool similar to Nano Banana, allowing image editing via text prompts. Its key differentiator is the ability to apply edits as a gradient, providing fine control over the strength or intensity of the edit.
Key Features:
- Controlled Intensity: Users can adjust a guidance slider to incrementally apply edits, observing the transformation.
- Scene Preservation: At lower strengths, Greg can apply edits while preserving the background and other elements of the scene, unlike many editors that alter the entire image.
- Integration: Can be used with other image editors like Step 1x Edit and Quinn Imageedit to improve their results by preserving original details and context.
- Availability: A GitHub repo is available, with plans for a Hugging Face space and code release.
Udio Music Generator: Licensing and Download Restrictions
Udio, a leading AI music generator, has announced a partnership with Universal Music Group (UMG) following a copyright infringement lawsuit. This partnership involves new licensing agreements that significantly impact users.
Consequences:
- Download Restrictions: Downloads from the platform are currently unavailable.
- Limited Download Window: A 48-hour window (starting Monday, November 3rd) has been provided for users to download existing songs. After this period, downloads will be permanently disabled.
- User Outrage: The sudden restrictions and lack of prior notice caused significant user backlash.
- Future Uncertainty: The tight control and regulation due to the partnership raise concerns about users' ability to create freely without copyright infringement issues, potentially marking the "death" of Udio as a user-friendly platform.
Minimax Music 2.0: AI Music Generation
Minimax Music 2.0 is a new AI music generator that allows users to create songs from text prompts, with the option to input custom lyrics or have lyrics generated automatically.
Features:
- Prompt-Based Generation: Describe the desired song, genre, and mood.
- Lyric Input/Generation: Users can provide their own lyrics or have them generated.
- Quality: Produces realistic vocals and consistent instruments and melodies.
- Genre Versatility: A single voice can be used to sing songs in various genres (e.g., blues, rock, electronic).
- Cost: Uses a credit system (300 credits per song), with new users receiving over 10,000 free credits.
Minimax Highaw 2.3 and M2
Minimax has also released Highaw 2.3, an AI video model noted for its superior physics and world understanding compared to models like V3.1 and Sora 2, particularly in high-action and physically complex scenes.
Furthermore, Minimax M2 has been open-sourced. It is currently considered the number one open-source AI model, with intelligence comparable to leading closed models like GPT-5, High, and Grok 4, and it outperforms Gemini 2.5 Pro. Model weights are available for local use, on-premise deployment, and fine-tuning.
Foley Control: AI Audio Generation for Silent Videos
Foley Control by Stability AI is a tool that generates audio for silent videos. This is particularly useful as most AI video models do not produce sound.
Capabilities:
- Sound Effect Generation: Detects actions in videos and generates synchronized sound effects.
- Timing Accuracy: Accurately times sound effects with on-screen actions, including footsteps.
- Input: Accepts silent video and a text prompt describing the desired audio or an audio sample.
- Architecture: Utilizes a frozen latent audio model with cross-attention bridges for synchronization.
- Availability: Currently only a technical paper and project page are available; no code has been released yet.
Conclusion
The past week has seen an explosion of advancements in AI, particularly in generative models for video, 3D worlds, and music, alongside significant progress in AI agents for gaming, security, and creative tasks. Key themes include the push for longer video generation (Long Cat Video), enhanced multimodal capabilities (Emu 3.5), efficient processing of vast amounts of data (Kimmy Linear), and novel approaches to generation that bypass traditional limitations (SVG). The development of more capable and versatile AI agents (Game Tars, Artvar, Emergent) signals a move towards more autonomous and integrated AI systems. While some tools offer impressive speed and efficiency (Nitro E, Pomelli), others are still grappling with quality and real-world autonomy (Neo1X robot). The landscape is also marked by evolving ethical considerations and licensing challenges, as seen with Udio's partnership with UMG. The rapid pace of innovation suggests a future where AI plays an increasingly integral role across various industries and daily life.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

Seedance 2.0 4K: The New AI Video King?
Zubair Trabzada | AI Workshop

New top local AI image generator is here! Already uncensored
AI Search

How to Make 4K AI Videos That Look REAL (Seedance 2.0 Full Guide) | Higgsfield Seedance 2.0 4k
ManuAGI - AutoGPT Tutorials

This AI Video Is 4K Now — and You CAN'T Tell It's AI | Higgsfield Seedance 4k
ManuAGI - AutoGPT Tutorials

The ONLY Place You Can Use Seedance UNLIMITED (No Credits, No Caps) | Higgsfield Seedance Unlimited
ManuAGI - AutoGPT Tutorials

Fable 5 COMING BACK! Deepseek v4.1, GPT-5.6 Leaks, Fusion API, & Kimi K2.7 Code High Speed! AI NEWS!
WorldofAI

"ChatGPT Moment" for Robotics Is Coming. The Real Problem Isn't Intelligence | Stanford, Catie Cuan
EO