3D waifus, AI for cancer, new AI image editors, new open-source king, Veo 3.1 updates - AI NEWS
By AI Search
Key Concepts
- DIT 360: AI model for generating high-quality panoramic images from text prompts or reference images.
- Puffin: Unified AI model that understands and generates images based on camera perspective and spatial information, combining vision, language, and camera data.
- Stream VLM: AI that analyzes videos in real-time, capable of narrating scenes and remembering past events within long videos.
- Dream Omni 2: Free and open-source image editor from ByteDance for image manipulation, including object replacement, lighting transfer, style transfer, and background replacement.
- Deep Somatic: AI tool for identifying cancer-related genetic mutations in tumors by analyzing DNA.
- UpToYou: AI method for reconstructing complete 3D models of people from multiple 2D photos, including texture.
- Uni Tree G1: Highly flexible and acrobatic humanoid robot demonstrating advanced kung fu routines and fluid movements.
- Ring 1T: One trillion parameter open-source thinking model from Ant Group, comparable to top closed-source models like Gemini 2.5 Pro and GPT-5.
- Fizz HSI: System enabling humanoid robots to interact naturally with the real world through tasks like carrying, sitting, and lying down.
- TAG: AI plugin to improve AI-generated image quality by reducing hallucinations.
- DGX Spark: Compact AI supercomputer from Nvidia for developers to prototype and run AI tools locally.
- Nano Banana: AI image editor integrated into Google Search for Android users.
- Veo 3.1: Google's latest video model with slight improvements in audio quality and character consistency.
- RTFM (Real-Time Frame Model): AI that creates interactive 3D-like video worlds in real-time.
- DTOE (Data-driven Training for Out-of-Distribution Embodiment): Framework using gaming data to train AI models for real-world robot control.
- MVP 4D: AI that creates animatable 3D heads from a single 2D photo.
AI Innovations and Open-Source Advancements This Week
This week has seen significant advancements in AI, with a notable surge in open-source contributions catching up to or even surpassing closed-source models. Key developments include new image generation and editing tools, sophisticated video analysis, advancements in humanoid robotics, and powerful new large language models.
DIT 360: Panoramic Image Generation
Insta360 has released DIT 360, an AI model capable of generating high-quality panoramic images from text prompts or reference images. Unlike existing image generators, DIT 360 excels at creating super long and panoramic scenes with high resolution and detail. It supports both outdoor and indoor scenes and offers inpainting (filling in blank areas) and outpainting (extending image edges) capabilities. A free Hugging Face space is available for online testing, generating 2048x1024 images in about a minute. The code is also released on GitHub, requiring a CUDA GPU but with unspecified VRAM requirements.
Puffin: Camera-Aware AI for Spatial Understanding
Puffin, developed by the team behind "Thinking with Camera," is a unified AI model that integrates vision, language, and camera data. It understands and generates images based on camera perspective and spatial information, allowing it to estimate camera parameters (roll, pitch, field of view) and generate images from specific camera angles. This capability enables the reconstruction of 3D maps from single images by generating multiple views. The model utilizes "camera tokens" for input, alongside text and images. A dataset of 4 million image, text, and camera data samples has also been released for training custom models.
Stream VLM: Real-Time Video Analysis
Stream VLM is an AI model designed for real-time understanding and analysis of videos, even long and complex ones. It can narrate sports footage with impressive accuracy, processing videos at up to eight frames per second on a single H100 GPU and outperforming GPT-4o Mini in benchmarks. Stream VLM employs a "reuse KV" technique for efficient information storage, enabling it to maintain memory of events within long videos. The code and dataset are publicly available on GitHub.
Dream Omni 2: ByteDance's Open-Source Image Editor
ByteDance has launched Dream Omni 2, a free and open-source image editor. This tool allows for sophisticated image manipulation, including:
- Object Replacement: Replacing elements in one image with objects from another.
- Lighting Transfer: Applying the lighting from one photo to another.
- Style Transfer: Applying the style of one photo to another, such as turning a realistic photo into a 3D pixel style or anime.
- Pose and Expression Transfer: Transferring the pose or expressions of one character to another.
- Pattern and Texture Transfer: Applying patterns or textures from one object to another.
- Background Replacement: Replacing the background of an image. It can merge up to four reference images. A free Hugging Face demo is available, and the code is on GitHub. The model size is approximately 16 gigabytes, suggesting it could run on 12GB of VRAM.
Deep Somatic: AI for Cancer Mutation Detection
Google has introduced Deep Somatic, an AI tool designed to identify cancer-related genetic mutations in tumors. It analyzes DNA from tumor and normal cells, converting genetic data into images that are then processed by a convolutional neural network. This CNN distinguishes between normal and cancerous cells, identifying inherited genetic variants, tumor-specific mutations, and sequencing errors. Trained on breast and lung cancer data, it has shown success in identifying mutations in aggressive brain cancer and pediatric leukemia, discovering previously known and 10 new variants with higher accuracy than existing tools. Both the tool and its training dataset are openly available.
UpToYou: 3D Model Reconstruction from Photos
UpToYou is an AI method that reconstructs complete 3D models of individuals, including textures, from multiple 2D photos taken at different angles and poses. It demonstrates high accuracy in geometry (15-18% more than other methods) and texture fidelity (21-46% more). A significant advantage is its speed, generating a full 3D model in just 1.5 minutes, compared to hours for methods like Dream Booth. The code is available on GitHub, with a Hugging Face folder indicating a total size of 3.2 GB, suggesting lower VRAM requirements.
Humanoid Robot Advancements
- Uni Tree G1: This highly flexible humanoid robot has showcased an upgraded kung fu routine with impressive speed, fluidity, and naturalness. It can perform continuous flips and complex routines, demonstrating significant improvements in balance and movement through reinforcement learning. Live demos and its shadow further validate its capabilities.
- Robotic Arm for Shaving: A demo from No Matrix featured a robotic arm designed to shave faces. However, the demo was noted for its hard cuts and lack of demonstration on tricky shaving areas, raising questions about its reliability.
- Fizz HSI: This system enables humanoid robots to interact naturally with the real world. The Uni Tree G1, using Fizz HSI, can autonomously carry objects of varying sizes and weights, sit down on chairs naturally, and lie down with human-like movements. The system trains robots in simulation using physically realistic movements compared to human motion data, then deploys them in the real world using sensors for accurate object localization and interaction. The code is publicly available.
Ring 1T: A Trillion-Parameter Open-Source Thinking Model
Alibaba's Ant Group, through its Inclusion AI team, has released Ring 1T, a one trillion parameter open-source "thinking" model. Based on their Ling 2.0 architecture and trained on the Ling 1T base foundation model with large-scale verified reward reinforcement learning, Ring 1T demonstrates performance comparable to or exceeding top closed-source models like Gemini 2.5 Pro and approaching GPT-5. It excels in benchmarks for competitive math, coding, and ARGI (learning new things on the fly). Notably, it achieved a silver medal at the International Math Olympiad, showcasing its deep thinking capabilities despite being a natural language model. The models are available on Hugging Face and ModelScope, with a total file size of approximately 2 TB, requiring significant computational resources.
TAG: Reducing Hallucinations in AI Images
TAG is an AI plugin designed to improve the quality of AI-generated images by reducing hallucinations. It amplifies the tangential component of the image generation process, steering the output towards more realistic and accurate representations without additional training or computational overhead. While demonstrated on older models like Stable Diffusion 1.5, 2.1, and SD3, the effectiveness on newer models like SD3.5 and advanced open-source editors like Quen Image or Hydream is questioned due to their already low hallucination rates. The code is available on GitHub.
Nvidia DGX Spark: Personal AI Supercomputer
Nvidia is rolling out DGX Spark, a compact AI supercomputer designed for developers. This desk-sized system features an Nvidia Grace Blackwell superchip and 128 GB of unified memory, supporting AI models up to 200 billion parameters. Multiple DGX Sparks can be linked for larger models. Priced at around $4,000, it targets serious AI researchers for secure local prototyping and validation, particularly in privacy-sensitive industries. Jensen Huang has personally delivered units to XAI and OpenAI.
Google Updates: Nano Banana in Search and Veo 3.1
- Nano Banana in Google Search: Google has integrated Nano Banana, a powerful AI image editor, into Google Search on Android devices. Users can access it via the lens mode, take a selfie, and prompt it to generate various images of themselves, such as a photo booth strip with an old-school look. This feature is currently available in English in the US and India.
- Veo 3.1: Google's latest video model, Veo 3.1, offers slight improvements in audio quality and character consistency over Veo 3.0. A new feature in Google Flow allows users to edit generated videos by inserting objects via prompts and drag-and-drop. The "ingredients to video" feature now supports uploading collages or grids of multiple reference objects as a single image.
RTFM: Real-Time Interactive 3D Video Worlds
World Labs has released RTFM (Real-Time Frame Model), an AI that creates interactive 3D-like video worlds in real-time. It runs efficiently on a single H100 GPU and uses a neural network trained on vast amounts of video data to learn lighting and physics, generating persistent scenes. Users can interact with these worlds, and the environment remains consistent even after leaving and returning. RTFM offers significantly improved quality compared to previous World Model releases and is available for online testing.
DTOE: Training Robots with Gaming Data
DTOE (Data-driven Training for Out-of-Distribution Embodiment) is a framework that uses gaming data to train AI models for real-world robot control. By studying gameplay scenes from various games (e.g., Broat, CS:GO 2, Minecraft), DTOE learns to predict player actions and generalize to new games. A vision-action foundation model then transfers these learned skills to physical robots. This approach has shown remarkable success, achieving a 96.6% success rate on robot manipulation benchmarks and 83.3% on navigation benchmarks, significantly outperforming previous models without direct real-world robot training. This method promises to make robot training cheaper and faster. The dataset and model are planned for release.
MVP 4D: 3D Heads from Single Photos
MVP 4D is an AI that generates animatable 3D heads from a single 2D photo. It uses a two-stage process involving a morphable multi-view video diffusion model to create multiple video angles of the face, which are then merged into a real-time renderable 4D representation. While it can create interactive 3D heads of celebrities like Morgan Freeman and Tom Cruise, the quality is noted as not always perfect, with some angles or expressions not being entirely accurate, and teeth rendering issues observed. The code is not yet released but is expected soon.
Conclusion
The AI landscape continues to evolve at an unprecedented pace, with open-source models increasingly challenging established closed-source leaders. This week's announcements highlight significant progress in image and video generation, real-time world rendering, sophisticated robotics, and the development of powerful, versatile AI models. The trend towards more accessible and efficient AI tools, from personal supercomputers to integrated search features, suggests a future where advanced AI capabilities are more widely available.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

How the hometown humiliation of Putin marks a turning point for Ukraine | DW News
DW News

Shocking video shows moment paramedics are hit by Israel in 'double-tap' strike
Sky News

Every Kind of Volcano | SciShow Kids
SciShow Kids

Putin Xi, To Catch a Castro, Red Carpet Rebellion • FRANCE 24 English
FRANCE 24 English

Pokemon goes prehistoric at Chicago's Field Museum
Reuters

Pokemon goes prehistoric at Chicago's Field Museum
Reuters

Trump's supporters furious over Trump smartphone scam.
ABC News In-depth