Infinite AI video, 4K images, realtime videos, DeepSeek breakthrough, Google’s quantum leap: AI NEWS

AI SearchAbout 13 min readOct 26, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • 3D World Generation: Creating three-dimensional environments from input data.
  • Multi-shot Video Generation: Producing videos composed of multiple distinct scenes or shots.
  • High-Resolution Image Generation: Creating images with exceptional detail and clarity.
  • Image Editing with Motion: Manipulating specific parts of an image by drawing and indicating movement.
  • Real-time Video Generation: Producing video content with minimal latency.
  • High-Resolution Video Generation: Creating videos with superior detail and clarity.
  • Video Editing with Text Prompts: Modifying existing videos using natural language instructions.
  • 3D Model Editing with Text Prompts: Altering 3D models through textual descriptions.
  • Agentic Systems for Video Generation: AI systems that autonomously refine video generation processes.
  • Humanoid Robots: Robots designed to resemble human form and movement.
  • Realistic Robot Faces: Advanced robotic facial features designed for lifelike expressions.
  • Long-form Video Generation: Creating videos of extended duration with consistent content.
  • Quantum Computing Breakthroughs: Significant advancements in the field of quantum computation.
  • Geospatial Reasoning: AI capabilities for analyzing and understanding Earth-related data.
  • AI-powered Web Browsers: Browsers integrated with AI assistants for enhanced functionality.
  • Visual Tokens for AI: Representing data as images or visual elements for AI processing.
  • Robot Motion Learning: Training robots to move safely and efficiently.

Tencent Hunyan World Mirror: Flexible Open-Source 3D World Generator

Tencent has released Hunyan World Mirror, described as the most flexible open-source 3D world generator currently available.

  • Functionality: It allows users to generate 3D worlds using an input image as a reference.
  • Input Flexibility: Beyond a single image, it can incorporate camera intrinsics, depth maps, poses, or any combination of these to guide the 3D world generation.
  • Output Capabilities: The tool can also generate camera parameters, depth maps, and normal estimations.
  • Examples:
    • Stitching together two input photos to create a unified 3D model.
    • Reconstructing a room from four input photos, estimating camera positions.
    • Generating detailed 3D point maps from multiple zoomed shots of a scene.
  • Output Formats: It can produce both 3D point maps (collections of scattered dots) and complete 3D scenes, including detailed reconstructions of objects like lighthouses with windows, fences, and flags.
  • Availability: The models are released on HuggingFace, and the GitHub repository provides instructions for local installation.
  • Requirements: A CUDA GPU is required. While specific VRAM requirements are not stated, the model size of 5 GB suggests compatibility with most consumer-grade GPUs.

Ant Group Hollow Scene: Text-Prompted Multi-Shot Video Generation

Ant Group has launched Hollow Scene, an open-source tool for generating entire multi-shot, long videos solely from text prompts.

  • Problem Addressed: Traditional video models are limited to generating short, isolated clips (5-10 seconds). Hollow Scene overcomes this by enabling the creation of longer, coherent videos.
  • Prompt Formatting:
    • A global caption defines the overall scene.
    • Characters are denoted by numbers (e.g., Character 1, Character 2).
    • The prompt specifies the number of shots and details each shot individually.
  • Example 1 (Sunrise): A prompt detailing a wide shot of a sunrise, followed by medium and close-up shots of a character observing it, resulting in a cinematic and consistent video.
  • Example 2 (Scientist and Android): A prompt describing a scientist and an android tending a garden of synthetic plants, with specific shots detailing their interactions and the environment, leading to a coherent and detailed output.
  • Capabilities: The AI excels at following detailed prompts, maintaining scene consistency, colors, and character appearance across shots.
  • Limitations: It is primarily suited for cinematic shots and may not perform as well with high-action or physically challenging scenes.
  • Availability: Models are released, and the GitHub repository provides installation instructions.
  • Technical Basis: It is based on Stable Diffusion 2.1 14B, utilizing the latest open-source version.
  • Generation Parameters: Prompts can include the number of frames, with a default of 10 seconds at 24 frames per second, allowing for longer video generation.

DYP: Open-Source AI for Native High-Resolution Image Generation

DYP is an open-source AI capable of generating extremely high-resolution images natively.

  • Key Feature: Generates images with exceptional detail, approaching 4K resolution.
  • Comparison with Flux: When compared to existing models like Flux, DYP significantly enhances detail and clarity, particularly at high resolutions. Examples show DYP producing sharper details in faces, grass, armor, and background elements.
  • Availability: The code is released on GitHub, with instructions for local installation.

In Paint for Drag: Image Editing by Drawing and Moving Objects

In Paint for Drag is an AI tool that allows users to edit images by painting over areas and indicating movement.

  • Process:
    1. Paint Over: Users paint over the specific areas of an image they wish to edit.
    2. Masking: The AI automatically selects or masks these painted areas.
    3. Draw Arrows: Users draw arrows to control the desired movement of the selected objects.
    4. AI Stitching: The AI moves the objects and seamlessly stitches the changes into the image.
  • Examples:
    • Moving a wing on a creature.
    • Repositioning a panda within an image.
    • Adjusting the position of an arm.
  • Comparison: Similar to Nano Banana, but instead of text prompts, it uses direct drawing and movement indication.
  • Availability: Code is released on GitHub, and a Colab demo is available for online testing.

Crea Realtime 14b: Open-Source Real-time Video Generation

Crea Realtime 14b is a new open-source video model that aims for real-time video generation.

  • Technical Basis: Based on Alibaba's Stable Diffusion 2.1 14B, enhanced with a new "self-forcing" technique.
  • Performance: Claims text-to-video inference at 11 frames per second on a single Nvidia B200 GPU.
  • Requirement: Requires a powerful GPU (like the B200) for real-time performance; not feasible on consumer GPUs.
  • Video-to-Video Capabilities: Can transform rough video compositions into detailed scenes, such as generating a race car scene from a basic layout.
  • Creative Potential: Enables transforming webcam footage into entirely different videos.
  • Model Size: The model is significantly larger than existing real-time video models, with a total size of nearly 60 GB, contributing to the high hardware requirements.
  • Availability: While not runnable on most consumer devices, cloud GPU rental is an option for testing.

Ultragen: High-Resolution (4K) Video Generation

Ultragen is an AI tool capable of generating high-resolution videos, including native 4K.

  • Key Feature: First model known to generate native 4K videos.
  • Quality: Produces more detailed and clearer videos compared to leading open-source models like Hunyan Video and Stable Diffusion 1.5. Examples highlight improved detail in mountains and buildings.
  • Speed Comparison:
    • Generating a 1080p video with Stable Diffusion 1.5 takes 35 minutes. Ultragen takes less than half that time.
    • Generating a 4K video with Stable Diffusion 1.5 is blurry and takes nearly 9 hours. Ultragen generates a high-quality 4K video in under 2 hours.
  • Prompt Examples: Demonstrates superior detail in a "snowy village at Dusk" prompt, with visible window details in Ultragen's output.
  • Performance Benchmarks: Ultragen outperforms leading video models in both 1080p and 4K benchmarks for quality.
  • Methodology: Employs a special attention mechanism that breaks down video generation into global models (overall scene) and local models (fine details), combining them for consistent and detailed scenes.
  • Availability: A technical paper and GitHub repository are released; code is expected soon.

Ditto: Text-Prompted Video Editing

Ditto is an AI video tool that allows editing existing videos using text prompts.

  • Editing Capabilities:
    • Completely change characters or backgrounds.
    • Insert objects into videos.
    • Transform anime scenes into realistic ones using a fine-tuned model.
  • Availability: Open-source and usable immediately.
  • Model Size: Each model is 6 GB, making it runnable on most consumer-grade GPUs.
  • Installation: A full installation tutorial is available.

3D Model Editor: Text-Prompted Micro-Editing of 3D Models

This AI tool enables micro-editing of existing 3D models using text prompts, similar to Nano Banana but for 3D.

  • Editing Capabilities:
    • Modify object attributes (e.g., "make the backpack bigger").
    • Replace objects (e.g., "replace the blue jacket with the brown down jacket").
    • Remove objects (e.g., "remove the chimney").
    • Add objects (e.g., "add a satellite dish on top of this house").
    • Modify object features (e.g., "get this dragon to hold a sword," "remove the wings of this dragon").
  • Consistency: Edits are applied locally without affecting the rest of the model.
  • Methodology: Combines FlowEdit (a 3D model editor) with Trellis (a 3D model generator) to enable local edits without manual masking.
  • Availability: Code is not yet released, but a dataset and Gradio demo are planned.

Google Vista: Agentic System for Automatic Video Generation Improvement

Vista is an agentic AI system designed to automatically improve video generation.

  • Process:
    1. Initial Generation: Takes a user prompt and generates multiple videos.
    2. Competition & Selection: Videos undergo rounds of evaluation to select the best.
    3. AI Critique: Specialized AI agents critique videos based on visual fidelity, motion, dynamics, audio, and context.
    4. User Feedback: Users can provide feedback on preferred videos.
    5. Reasoning Agent: A reasoning agent rewrites and improves the original prompt based on critiques and feedback.
    6. Iterative Refinement: The cycle repeats, with each loop generating a better-looking video.
  • Examples:
    • "Spaceship entering hyperdrive" prompt improved by Vista to include a more cinematic tunnel effect.
    • "Block of ice melting into a drink" prompt enhanced with more detailed descriptions by Vista.
    • "Couple running through a downpour" prompt significantly improved for coherence and realism.
    • "Aerial view of a lush green forest" prompt enhanced with cinematic camera motions and better colors/sound.
  • Availability: Currently, only a technical paper is released.

Chat LLM by Abacus AI: All-in-One AI Platform

Chat LLM by Abacus AI is a sponsored platform offering integrated access to various AI models.

  • Features:
    • Seamless switching between different AI models.
    • Access to top image and video generators.
    • Artifacts Feature: Previewing generations side-by-side for coding.
    • Deep Agent Feature: Autonomous execution of complex tasks like creating presentations, websites, and research reports.
  • Pricing: $10 per month for access to all features.

Unree Robotics H2: Advanced Humanoid Robot

Unree Robotics has released the H2, an advanced humanoid robot.

  • Comparison: An evolution from their previous G1 model, known for acrobatics.
  • Movement: Exhibits flexible, natural, and human-like movements, including dancing, kung fu, and ballet spins.
  • Specifications:
    • Height: Approximately 180 cm.
    • Weight: Around 70 kg.
    • Degrees of Freedom: 31, including 3 for the waist and 2 for the neck.
    • Arms: 7 degrees of freedom each for dextrous manipulation.
  • Facial Features: Features a "bionic human face" intended to be lifelike and relatable, though described as "creepy" by the presenter.

Head AI Origin M1: Realistic Robot Face

Head AI Origin M1 by a Shanghai-based company features a highly realistic male robot face.

  • Design: Focuses on ultra-realistic facial expressions.
  • Actuation: Equipped with up to 25 brushless micro motors beneath synthetic skin for subtle movements like blinking, eyebrow raising, and head tilting.
  • Sensors: Embedded cameras in the eyes for gaze tracking and environmental perception.
  • Interaction: Recognizes and understands speech, responds in real-time with lip-syncing.
  • Presenter's Preference: The presenter expresses a preference for more stylized faces (e.g., catgirl).

Stable Video Infinity: Infinite Length Video Generation

Stable Video Infinity is an AI tool capable of creating videos of infinite length with consistent scenes and smooth transitions.

  • Key Feature: Generates long videos (over 40 seconds, claimed up to 10 minutes) without scene deformation or errors, unlike competitors.
  • Consistency: Maintains visual coherence throughout extended video generations.
  • Audio Synchronization: Can lip-sync to provided audio.
  • Example (Obama): Demonstrates consistent lip-syncing for over a minute, with Obama's face remaining stable, unlike other tools.
  • Availability: Code is released with instructions for local execution.
  • Hardware: Used an A100 with 80 GB for development, but the model is based on Stable Diffusion 2.1 14B, suggesting potential for consumer GPU compatibility.

Google Quantum Computing Breakthrough: Willow Chip and Quantum Echoes

Google has announced a significant breakthrough in quantum computing with its Willow quantum chip.

  • Quantum Computing Basics: Contrasts regular computers (binary 0 or 1) with quantum computers (dimmer switches, continuous values), enabling greater computational power.
  • Challenge: Quantum computers are sensitive and prone to errors due to their continuous nature.
  • Willow Chip: An advanced quantum computer that successfully ran a complex quantum algorithm.
  • Quantum Echoes: A new technique where a signal is sent into a quantum system, reversed, and the resulting echo is measured to understand how disturbances spread. This method is precise and reliable.
  • Performance: The Willow chip ran the quantum echoes algorithm 13,000 times faster than classical supercomputers.
  • Impact: This precision allows for better study of the universe, from molecules to black holes, fundamental to chemistry, biology, and material science, potentially leading to breakthroughs in biotechnology, solar energy, and nuclear fusion.
  • Availability: A technical paper is available.

Google Earth AI: Geospatial Reasoning with Gemini

Google has updated Google Earth AI with advanced reasoning capabilities powered by Gemini.

  • Functionality: Combines decades of Earth modeling data with Gemini's reasoning abilities to answer complex questions about Earth data autonomously.
  • Geospatial Reasoning Framework: Links various Earth data models (weather, population, satellite imagery) to identify vulnerable communities, disaster risks, deforestation, and climate change impacts.
  • User Interaction: Users can prompt the AI with queries, and it will automatically find and analyze relevant information in satellite imagery, eliminating the need for manual searching.
  • Target Audience: Initially available in the US for Google Earth Professional and Professional Advanced users, and for Google AI Pro and Ultra subscribers. Primarily targeted towards enterprises and research labs.
  • Future: AI-powered Google Earth and Maps are expected in the near future.

OpenAI ChatGPT Atlas: AI-Powered Web Browser

OpenAI has launched ChatGPT Atlas, a web browser powered by ChatGPT.

  • Architecture: A Chromium wrapper, essentially a clone of Google Chrome with an integrated ChatGPT sidebar.
  • Functionality:
    • Ask ChatGPT questions about the content of the current webpage.
    • Refine text by highlighting and using a ChatGPT button.
    • Browser Memories (Optional): Remembers browsing history and context to recall previously viewed items.
    • Agent Mode (Preview for Plus/Pro/Business users): Allows ChatGPT to perform autonomous tasks on the browser, such as shopping (adding items to cart and checking out).
  • Comparison: Similar to Perplexity's Comet browser and earlier Chrome extensions.
  • Availability: Currently available for macOS for free; Windows, iOS, and Android versions are coming soon. Agent mode is in preview.

DeepSeek OCR: Visual Tokens for AI Model Design

DeepSeek has published a paper on DeepSeek OCR, proposing a paradigm shift in AI model design using visual tokens.

  • Problem: Traditional LLMs convert text to numerical tokens, which becomes computationally expensive for long texts.
  • Solution: Instead of text tokens, DeepSeek OCR uses visual tokens by taking screenshots of text pages and processing them as images.
  • Architecture:
    1. Screenshot Input: Takes an image of the text.
    2. Chunking: Breaks the image into smaller chunks.
    3. Deep Encoder:
      • SAM (Segment Anything Model): Focuses on local areas for fine details (local attention).
      • Downsampling: Compresses data via a convolutional neural network.
      • CLIP Component: Understands context and relationships across the entire image (global attention).
    4. DeepSeek 3B Model: A small, efficient model (570 million active parameters) decodes vision tokens back into text.
  • Compression and Accuracy:
    • 10x compression ratio (1000 text tokens to 100 vision tokens) achieves ~97% decoding precision (nearly lossless).
    • 20x compression ratio (1000 text tokens to 50 vision tokens) achieves ~60% precision, a trade-off between efficiency and accuracy.
  • Advantages: Can analyze tables, diagrams, charts, and other non-textual elements. Supports multilingual recognition for nearly 100 languages.
  • Potential Impact: May lead to a shift towards vision tokens as the primary input for future AI models.
  • Availability: Code is released on GitHub.

Soft Mimic: Safe and Smooth Robot Motion Learning

Soft Mimic is an AI system that helps robots learn to move more safely and smoothly, mimicking human motion.

  • Problem: Traditional motion tracking methods for robots are stiff, leading to awkward and error-prone movements (e.g., robots failing to balance or knocking over objects).
  • Methodology:
    1. Human Motion Input: Takes video of human movement.
    2. Inverse Kinematic Solver: Helps the robot learn safe and compliant movements.
    3. Improved Dataset: Creates a dataset of all possible movements.
    4. Reinforcement Learning: Trains robots in simulation using this dataset over tens of thousands of iterations.
    5. Stiffness Control: Allows adjustment of movement stiffness.
  • Benefits: Robots become more adaptable, flexible, and natural in their movements.
  • Examples:
    • Robots can move flexibly while maintaining balance, even when pushed.
    • Robots can pick up boxes of different dimensions without issue.
    • Robots can navigate obstacles and pour liquids without spilling, especially with lower stiffness settings.
  • Availability: Only a technical paper is released; code release is uncertain.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.