New AI beats NanoBanana, video to 3D, AI Minecraft, dish-washing robots, new open source models

AI SearchAbout 9 min readSep 7, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • 3D World Generation: Creating interactive 3D environments from images or videos.
  • 3D Model Reconstruction: Generating 3D models from multiple images or videos of an object.
  • Text-to-Speech (TTS) and Voice Cloning: Synthesizing speech from text and replicating a person's voice.
  • Image Editing and Generation: Modifying existing images or creating new ones from scratch using AI.
  • Open-Source AI Models: AI models with publicly available code and data, allowing for modification and distribution.
  • Large Language Models (LLMs): AI models with billions of parameters, capable of complex tasks like coding and research.
  • Minecraft Simulation: Creating interactive environments similar to Minecraft using AI.
  • Mixture of Experts (MoE): An AI architecture that combines multiple specialized models to improve performance.
  • Real-time Generation: Generating content on the fly as the user interacts with the system.

Tencent Hunyuan World Voyager: 3D World Generation from a Single Image

  • Main Topic: Tencent's Hunyuan World Voyager generates consistent 3D point clouds and navigable 3D worlds from a single reference image.
  • Key Points:
    • Unlike other game generators (Google's Genie 3, Tencent Hunyuan Gamecraft) that produce videos, Voyager creates a consistent 3D point cloud.
    • It jointly generates depth and RGB videos, enabling robust and direct 3D reconstruction.
    • Users can customize the style of the 3D scene (e.g., snowy scene, day/night).
  • Examples: Generating 3D worlds from images of interiors, demonstrating the preservation of details like paintings and furniture.
  • Comparison: Voyager preserves details from the original image more accurately than competitors, especially in areas outlined in red in the video.
  • Benchmarks: Voyager outperforms competitors in novel view synthesis, camera control, object control, content alignment, and 3D consistency.
  • Availability: Open-source on GitHub with instructions for local download and execution.
  • Technical Details: Requires a minimum GPU memory of 60 GB for 540p generation and 80 GB recommended.
  • Future: Potential for quantized and compressed versions to run on lower VRAM.

Recon Via Gen: 3D Model Reconstruction from Multiple Images or Video

  • Main Topic: Recon Via Gen creates accurate 3D models from multiple images or a video of an object.
  • Key Points:
    • Captures details from multiple angles (front, side, back), unlike single-image generators.
    • Can reconstruct 3D models from videos, even with complex scenes and shaky camera movement.
  • Examples: Generating 3D models of figures, characters, and scenes from videos.
  • Comparison: Outperforms Hunyuan 3D 2.5 multiview in capturing details and generating accurate 3D models, especially in complex scenarios (e.g., a car with six wheels, a penguin behind a plushy).
  • Availability: Free Hugging Face space for online testing and a GitHub repository with plans to release the complete model.
  • Process: Users upload a video or multiple images, adjust generation settings, and click "generate" to create a 3D model.
  • Example Usage: Uploading two images of a Mecca robot to generate a 3D model with surface details.

Tencent Audio Story: Generating Narrative Audio from Video

  • Main Topic: Tencent's Audio Story generates long-form narrative audio from a silent video input.
  • Key Points:
    • Trained on audio from Tom and Jerry episodes to apply appropriate audio to videos.
    • Audio is aligned with the actions in the video.
    • Can be applied to different animations (e.g., Donald Duck).
  • Limitations: Audio quality is not perfect, and the model is trained only on Tom and Jerry episodes, limiting the style of audio generated.
  • Availability: GitHub repository with instructions for local download and execution under the Apache 2 license (allowing commercial use).
  • Future: Plans to release the dataset and training code.
  • Technical Details: Requires an Nvidia CUDA GPU, but VRAM requirements are not specified.

Microsoft Vibe Voice Drama: Text-to-Speech and Voice Cloning

  • Main Topic: Microsoft's Vibe Voice, a text-to-speech and voice cloning tool, was taken down shortly after release.
  • Key Points:
    • Vibe Voice allowed cloning anyone's voice with just a few seconds of audio and generating outputs up to 90 minutes long with multiple speakers.
    • Two variants were released: a small 1.5 billion parameter model and a larger 7 billion parameter model (better quality).
    • Microsoft took down the 7 billion parameter model and the GitHub repository.
  • Speculation: Microsoft likely deemed the AI too dangerous due to its ability to clone voices accurately and generate potentially harmful content.
  • Workaround: The larger 7 billion parameter version is still available on ModelScope (Chinese version of Hugging Face).
  • License: Vibe Voice was released under the MIT license, allowing anyone to clone and re-upload it.
  • ComfyUI Integration: Instructions for installing Vibe Voice on ComfyUI are available, with updated URLs for downloading the large model.

Highaw O2: AI Video Generator with Start and End Frames

  • Main Topic: Highaw O2 is a state-of-the-art video model that excels in prompt understanding, physics, camera control, and overall coherence.
  • Key Features:
    • Start and end frames feature: Users upload reference images to use as the start and end frame, and Highaw seamlessly interpolates everything in between.
    • Camera movements can be inserted to make generations more cinematic.
  • Examples:
    • Time-lapse from day to night using two different images as start and end frames.
    • Metamorphosis time-lapse of a butterfly emerging from its cocoon.
    • Flower bud blooming using two different images as start and end frames.

Figure O2 Robot: Autonomous Dishwashing Demo

  • Main Topic: A new demo of the Figure O2 humanoid robot autonomously loading dishes into a dishwasher.
  • Key Points:
    • The robot can grasp and hold dishes and cups firmly, despite the challenges of robotic hands.
    • It can manipulate items of different shapes and sizes and successfully place them into the dishwasher.
  • Limitations:
    • The dishes were already on the table, and the robot didn't demonstrate grabbing them from elsewhere.
    • The demo didn't show the complete task, such as pushing the tray, closing the door, and starting the machine.

Chatterbox Multilingual: Open-Source Text-to-Speech and Voice Cloning

  • Main Topic: Chatterbox Multilingual is an open-source text-to-speech model that supports 23 languages and offers emotion exaggeration control and zero-shot voice cloning.
  • Key Features:
    • Supports 23 languages.
    • Emotion exaggeration control.
    • Zero-shot voice cloning: Clones voices accurately with just a few seconds of reference audio.
  • Examples:
    • Expressive control: Generating angry or emotional speech.
    • Multilingual support: Generating speech in Chinese, Japanese, Korean, Spanish, French, German, Italian, Russian, Arabic, and Hindi.
  • Availability: Free online demo on Hugging Face and a GitHub repository with instructions for local download and execution.
  • Technical Details: The model is small (2 GB), so it can run on GPUs with over 2 GB of VRAM.

DH3: Mysterious Image Editor Potentially Better Than Nano Banana

  • Main Topic: A new image editor called DH3 is being tested anonymously and may be better than Google's Nano Banana.
  • Key Points:
    • DH3 is being tested on the Image Arena by Artificial Analysis, where users can AB test different image editors.
    • DH3 has shown promising results in image editing and image generation tasks, often beating GPT-4o.
  • Limitations: Limited information is available about DH3, including its origin and capabilities.
  • Testing: Users can participate in AB testing on the Image Arena to compare DH3 with other models.

Kimik K2 0905: Upgraded Open-Source Coding Model

  • Main Topic: Kimik K2, an open-source coding model, has been upgraded with enhanced coding capabilities and a larger context window.
  • Key Improvements:
    • Enhanced aentic coding and improved front-end coding capabilities.
    • Increased context window from 128K to 256K tokens.
  • Benchmarks: Kimik K2 performs significantly better than the previous version in coding benchmarks and is as good as or better than Claude Sonnet 4.
  • Cost: Kimik K2 is completely free and open-source, making it much cheaper than proprietary models like Claude Sonnet 4.
  • Availability: Available for free use on kimmy.com and models are released on HuggingFace.
  • Example: Creating a beautiful CRM dashboard with real-time insights using HTML.
  • Comparison: Kimik K2's dashboard generation is more interactive and complete than Claude Opus 4.1.

Oasis 2.0: Real-Time Minecraft Generator

  • Main Topic: Oasis 2.0 by Decart is a real-time Minecraft generator that creates interactive environments on the fly.
  • Key Features:
    • Runs at 1080p resolution at 30 frames per second, allowing for seamless real-time interactivity.
    • Includes a video-to-video feature where users can insert prompts to change the style of the scene.
  • Limitations:
    • The online demo has issues with controls and disconnects after a few seconds.
    • The quality of the generation is not as detailed as other real-time game generators.
  • Availability: Online demo and model release for download and installation.
  • Technical Details: VRAM requirements are not specified, but a significant amount is likely needed for real-time generation.

Long Cat Flash: State-of-the-Art Open-Source Model by a Food Delivery Company

  • Main Topic: Long Cat Flash, a 560 billion parameter mixture of experts model developed by a Chinese food delivery company (Mtoan), is state-of-the-art in coding and instruction following.
  • Key Features:
    • Uses a shortcut connected mixture of experts structure for improved training and inference efficiency.
    • Has a dynamic computation mechanism that activates only a subset of the parameters needed for each task.
  • Benchmarks: Long Cat Flash is on par with leading open-source models (Deepseek, Quen, Kimik K2) and proprietary models (GPT-4o, Claude 4, Gemini 2.5 Flash) in various benchmarks.
    • Scores the highest compared to all other models on aentic tool use benchmarks.
  • Availability: Free online chat interface and models released on Hugging Face.
  • Technical Details: The model is massive (over 500 billion parameters), requiring significant GPU resources.
  • Examples:
    • Creating a beautiful CRM dashboard with HTML.
    • Analyzing e-commerce growth in Asia from 2020 to 2025 using web search.
    • Researching rehab protocols and return-to-sport timelines for a 25-year-old athlete with an ACL injury.

Conclusion

This week in AI has seen significant advancements across various domains, including 3D world generation, 3D model reconstruction, text-to-speech, image editing, and coding. Notably, open-source models are rapidly catching up to and even surpassing proprietary models in performance, as demonstrated by Kimik K2 and Long Cat Flash. The release of tools like Hunyuan World Voyager and Recon Via Gen enables users to create immersive 3D experiences and models from simple inputs. While some AI tools face challenges related to safety and ethical concerns, the overall trend indicates a democratization of AI technology and increased accessibility for developers and users alike. The advancements in humanoid robotics, as seen with the Figure O2 robot, showcase the potential for AI to automate everyday tasks.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.