New DeepSeek, image to video game, Google kills photoshop, robot army, 3D cloud to mesh. AI NEWS

AI SearchAbout 9 min readAug 25, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Video Lip Sync (Infinite Talk, MultiTalk)
  • Multimodal Models (Ren EC)
  • 3D Scene Editing (Tinker)
  • Real-time Video Game Generation (Mirage 2, Genie 3)
  • AI-Powered Image Editing (Google Photos, Nano Banana, Quinn Image Edit)
  • Large Language Models (LLMs) (Deepseek V3.1, GPT5, Quen)
  • 3D Mesh Generation from Point Clouds (Mesh Coder)
  • Humanoid Robotics (Boston Dynamics Atlas, AGI Bot A2)
  • 3D Object Segmentation (Geio SAM 2)

Infinite Talk: Advanced Video Lip Sync Tool

  • Main Topic: Introduction of Infinite Talk, a new video lip sync tool by Magen, the creators of MultiTalk.
  • Key Points:
    • Unlike other lip sync tools that use a static image, Infinite Talk uses a reference video to animate expressions and movements.
    • This video-to-video approach generates more realistic and natural results.
    • Allows control over camera movements and scene transitions, focusing on lip movements and expressions.
    • Also supports image-to-video, animating an image based on an audio clip.
    • Uses "long sequence sparse frame video dubbing" to maintain reference frames and interpolate between them while lip-syncing.
  • Examples:
    • Lip-syncing a video of a person talking about a phone review with a different audio track.
    • Animating singing examples with natural movements.
    • Changing the audio of a movie clip while maintaining realistic expressions.
  • Technical Details:
    • Uses one 2.1 as the base model.
    • Compatible with acceleration steps like Fusion EX or Liteex 2V.
    • Supports Comfy UI.
  • Availability: Code and technical report available on GitHub.

Ren EC: Alibaba's Multimodal Model

  • Main Topic: Alibaba's new multimodal model, Ren EC, which understands and interacts with the world through videos and images.
  • Key Points:
    • Can answer questions about objects and actions in videos.
    • Performs tasks like object segmentation and spatial understanding.
    • Predicts distances between objects and the camera.
    • Demonstrates strong spatial understanding, even with shaky videos and brief object appearances.
  • Examples:
    • Identifying the surface and function of an object.
    • Segmenting objects based on a question (e.g., "Which object should I use to drink water?").
    • Predicting the distance between two objects (e.g., 1.23m).
    • Determining which object is taller or above another.
    • Predicting the distance between the camera and an object (e.g., 1.63m).
    • Answering directional questions (e.g., "Is this object on your right rear or right front?").
  • Performance: Outperforms other multimodal and vision models in object and spatial cognition benchmarks.
  • Technical Details:
    • Uses Quen 2.5 as the base model.
    • Models have 2 billion or 7 billion parameters.
    • Operates under the Apache 2 license.
  • Availability: Open-sourced with instructions for local download and training.

Tinker: 3D Scene Editing with Prompts

  • Main Topic: Introduction of Tinker, an AI tool that allows editing of 3D scenes using text prompts.
  • Key Points:
    • Enables adding elements (e.g., autumn leaves, frost), changing textures, and altering the overall style of a 3D scene.
    • Can perform micro-edits, such as changing the color of an object or altering clothing.
    • Uses flux context, an AI image editor, to edit frames and then fills in the blanks to generate the rest of the 3D scene.
  • Examples:
    • Adding autumn leaves or a frost effect to a scene.
    • Changing a road to a river.
    • Transforming a scene into a Van Go style painting.
    • Changing the color of slippers to white.
    • Putting a man in a black tuxedo.
  • Performance: Outperforms other 3D scene editors like DGE or Edit Splat in terms of accuracy and consistency.
  • Availability: GitHub repo available with plans to release data, source code, and pipeline.

Mirage 2: Real-time Interactive Video Game Generator

  • Main Topic: Introduction of Mirage 2, a real-time video game generator that creates playable environments from images.
  • Key Points:
    • Generates playable 3D worlds from uploaded images in various styles (e.g., cyberpunk city, medieval town).
    • Allows users to chat with the AI to edit the scene and introduce new elements.
    • Supports real-time interaction with controls for movement, running, jumping, and attacking.
    • Can generate both first-person and third-person scenes.
    • Seamlessly transitions between different scenes based on prompts.
  • Examples:
    • Generating a 3D world from a simple drawing.
    • Transforming a kid's drawing into a playable environment.
    • Turning the Starry Night painting into a 3D world.
    • Seamlessly transitioning from a cyberpunk city to a rainforest, a mountaintop castle, or a tropical island.
  • Comparison: Compared to Google DeepMind's Genie 3, Mirage 2 is more interactive, while Genie 3 has higher quality and better physics understanding.
  • Technical Details:
    • Claims to run on a single consumer-grade GPU.
  • Availability: Demo available online.

ZAI: AI-Powered Platform for Content Creation (Sponsored)

  • Main Topic: Promotion of ZAI, an AI platform with state-of-the-art models like GLM4.5.
  • Key Points:
    • Offers a free online platform for using its AI models.
    • Features an AI slides tool for creating presentations.
    • Excels at coding, generating playable games from prompts.
    • Includes web search for fetching the latest information.
    • Offers a full-stack feature for creating complete apps with multiple pages.
    • Has a vision model for analyzing images.
  • Examples:
    • Creating a presentation on the wildlife of Sri Lanka.
    • Generating a space shooter game with one prompt.
    • Creating a financial report on Nvidia.
    • Designing a futuristic blog and forum with a landing page, topic hub, and blog posts.
    • Identifying the location of a photo (e.g., English Bay).

Google Photos AI Image Editing

  • Main Topic: Introduction of a new AI-powered image editing feature in Google Photos.
  • Key Points:
    • Allows users to edit images using text prompts.
    • Can remove glare, brighten photos, add clouds, remove objects, restore old photos, colorize images, swap clothes, change backgrounds, and adjust brightness, contrast, and saturation.
    • Eliminates the need for manual selection, masking, or adjusting sliders.
    • Potentially uses Google's stealth image editor model, Nano Banana.
  • Availability: Initially coming to the latest Google Pixel 10 in the US, then rolling out to other Android and iOS devices.

Deepseek V3.1: Enhanced Language Model

  • Main Topic: Release of Deepseek V3.1, the latest language model from Deepseek.
  • Key Points:
    • Offers both a "thinking mode" for complex reasoning and a "non-thinking mode" for faster responses to simpler tasks.
    • Runs faster than Deepseek R1.
    • Demonstrates stronger agentic skills.
    • Shows significant improvements in coding benchmarks compared to previous versions.
  • Examples:
    • Creating a space shooter game with asteroid fields and alien invaders.
    • Generating a crypto portfolio dashboard with price fluctuations, risk assessments, and trade simulators.
    • Researching surgical options and long-term prognosis for a newborn with congenital heart defects and compiling a comprehensive report.
    • Creating a visual simulation of a beehive construction. (Failed to fully execute the prompt).
  • Performance: Ranks highly on independent leaderboards, but its exact position varies depending on the benchmark.
  • Technical Details:
    • Pricing is cheaper than Quen 3.
  • Availability: Available for free on deepseek.com and for download on Hugging Face.

Mesh Coder: Point Cloud to Editable 3D Mesh

  • Main Topic: Introduction of Mesh Coder, an AI tool that converts point clouds into editable 3D meshes.
  • Key Points:
    • Outputs 3D models in the form of code that can be used in 3D editing software like Blender.
    • Allows users to easily edit the size, shape, and dimensions of any part of the 3D model.
    • Enables changing the depth of objects or increasing the resolution of the mesh.
    • Facilitates understanding the structure of the object by plugging the code through an LLM.
  • Examples:
    • Converting a point cloud of a sofa into an editable 3D mesh with code for each section.
    • Converting a point cloud of a chair into editable code.
    • Converting a point cloud of a toilet into editable code.
  • Availability: GitHub repo available with plans to release the code by November.

Humanoid Robotics Updates

  • Main Topic: Updates on humanoid robots, including Boston Dynamics' Atlas and AGI Robotics' AGI Bot A2.
  • Boston Dynamics Atlas:
    • Demonstrates autonomous object manipulation, picking up objects and placing them in a container.
    • Adapts to obstacles and changes in the environment.
    • Features interchangeable claw-like hands.
  • AGI Robotics AGI Bot A2:
    • Mass-produced in China, with around a thousand units already produced and plans to ramp up to 5,000 by the end of the year.
    • Deployed in commercial service environments for customer guidance, marketing presentations, and factory automation.
    • Features 40 active degrees of freedom, multiple sensors, and the ability to carry up to 15 kg.
  • World Humanoid Robot Games:
    • Updates from the games in Beijing, including a solo dance contest won by the Uni Tree G1.

Quinn Image Edit and Nano Banana: State-of-the-Art Image Editors

  • Main Topic: Introduction of Quinn Image Edit and Nano Banana, two state-of-the-art image editors.
  • Quinn Image Edit:
    • Allows users to edit images with prompts, without manual selection or masking.
    • Can change brightness, contrast, white balance, restore photos, change text, and micro-edit features.
    • Ranks as the best open-source image editor on the artificial analysis leaderboard.
  • Nano Banana:
    • A stealth model, rumored to belong to Google, considered the best AI image editor available.
    • Accessible on the LM Arena platform for blind testing.
    • Can restore and colorize damaged photos.
  • Availability: Quinn Image Edit has a full installation tutorial and review available. Nano Banana is only accessible on LM Arena.

Geio SAM 2: Accurate 3D Object Segmentation

  • Main Topic: Introduction of Geio SAM 2, a tool for accurate identification and segmentation of 3D objects.
  • Key Points:
    • Takes a 3D mesh as input and allows users to guide the segmentation with prompts, clicks, or boxes.
    • Generates 12 different views of the 3D object to improve segmentation accuracy.
    • Offers flexible and customizable segmentation based on user input.
  • Examples:
    • Segmenting a character in different ways, separating or merging parts of the body.
    • Accurately segmenting fruits from a bowl, unlike other segmentation models.
    • Accurately segmenting a well based on meaningful parts.
  • Availability: GitHub button available, but no code or models have been released yet.

Conclusion

This week in AI has seen significant advancements across various domains, including video lip sync, multimodal models, 3D scene editing, real-time video game generation, image editing, language models, 3D mesh generation, robotics, and 3D object segmentation. Tools like Infinite Talk, Ren EC, Tinker, Mirage 2, Google Photos' AI editing, Deepseek V3.1, Mesh Coder, Atlas, AGI Bot A2, Quinn Image Edit, Nano Banana, and Geio SAM 2 showcase the rapid progress and increasing capabilities of AI in creating, understanding, and manipulating digital content and physical environments. The open-sourcing of many of these tools and models further democratizes access to AI technology and fosters innovation.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.