New Deepseek, new top AI video & image models, Gemini 3 Deep Think, realtime TTS: AI NEWS
By AI Search
Key Concepts
- Text-to-Speech (TTS): Generating human-like speech from text.
- Voice Cloning: Creating a synthetic voice that mimics a specific person's voice.
- Real-time Generation: Producing output (speech, video, etc.) with minimal delay, suitable for live interaction.
- Parameter Count: A measure of a model's size and complexity, often correlating with its capabilities.
- GPU (Graphics Processing Unit): Specialized hardware for parallel processing, crucial for AI model execution.
- CPU (Central Processing Unit): The main processor of a computer.
- Speaker Similarity Score: A metric evaluating how closely a generated voice matches a reference voice.
- Open-Source Model: AI models whose code and weights are publicly available for use, modification, and distribution.
- Closed-Source/Proprietary Model: AI models whose code and weights are not publicly available.
- Video Generation: Creating video content from text prompts or other inputs.
- Binaural Audio/3D Audio: Audio designed to create an immersive, three-dimensional listening experience, often using headphones.
- Dual Branch Audio Generation: A technique for generating audio for left and right ears separately to enhance realism.
- Conditional Space-Time Module: A component in AI models that ensures generated content aligns with temporal and spatial actions in the input.
- Omnimodal Model: An AI model capable of understanding and processing multiple types of data (text, images, video).
- Distilled Model: A smaller, more efficient version of a larger AI model, often achieved through knowledge distillation.
- Humanoid Robot: Robots designed to resemble the human body in form and movement.
- Image Generation: Creating images from text prompts.
- Image Editing: Modifying existing images based on instructions.
- Parameter Size: The number of adjustable weights in a neural network, indicating its complexity.
- Agentic Operating System: An operating system that uses AI agents to autonomously execute complex workflows.
- Low Latency: Minimal delay in processing and output.
- High-Fidelity Emotion Capture: The ability to accurately represent subtle human emotions in generated content.
- Uncanny Valley: A phenomenon where AI-generated content that is almost, but not perfectly, human-like can evoke feelings of unease or revulsion.
- Distribution Matching Distillation: A technique for improving the efficiency and speed of generative models.
- Timestep Forcing Pipeline Parallelism (TPP): A method for parallelizing video processing across multiple devices to speed up generation.
- FPS (Frames Per Second): A measure of video playback speed.
- VRAM (Video Random Access Memory): Memory on a graphics card, crucial for running AI models.
- Poster Copilot: An AI agent designed to assist in creating professional posters and graphics.
- Multi-round Refinement: The ability to iteratively improve generated content through multiple prompts.
- Aspect Ratio: The proportional relationship between the width and height of an image or video.
- Compute: Computational resources required to run AI models.
- Reasoning Capacity: The ability of an AI model to perform complex logical deductions and problem-solving.
- Visual Puzzles: Challenges that require understanding and applying visual patterns.
- Mixture of Experts (MoE) Model: A type of neural network architecture that combines multiple specialized "expert" networks.
- Apache 2.0 License: A permissive open-source software license.
- Leaderboard: A ranking of AI models based on performance on specific benchmarks.
- Token: A unit of text or data processed by an AI model.
- Hugging Face: A platform and community for machine learning, hosting models and datasets.
- Depth Estimation: Predicting the distance of objects in an image from the camera.
- Normal Estimation: Determining the orientation of surfaces in an image.
- Core Predictor: The initial stage of a model that makes a rough prediction.
- Detail Sharpener: A component that refines predictions for greater accuracy and detail.
- Unified Model: A single AI model capable of handling multiple modalities (text, image, video).
AI Weekly Roundup: Explosive Releases in Video, Image, and Language Models
This week has seen an unprecedented surge in AI advancements, with the release of four new video generators, three state-of-the-art image generators and editors, and significant improvements in real-time text-to-speech and open-source language models. Notable developments include a real-time, infinitely long video generator, a superior animation tool, and a powerful open-source model rivaling top closed-source alternatives. Humanoid robot demonstrations and advanced AI models for complex reasoning also made headlines.
1. Vibe Voice: Real-time, Consumer-Grade Text-to-Speech
Main Topics and Key Points:
- Vibe Voice has released a real-time version of its text-to-speech (TTS) generator.
- It supports various accents and languages and can instantly clone voices with minimal reference audio (seconds).
- The real-time interface generates speech within milliseconds of pressing start.
- Technical Details: The model has only 0.5 billion parameters, is approximately 2 GB in size, and can run on consumer-grade GPUs or even CPUs. Speech generation takes around 300 milliseconds, achieving near real-time performance.
- Performance: It boasts a low error rate and the highest speaker similarity score compared to other models.
- Availability: The code is publicly available on GitHub with instructions for download and execution.
2. Steady Dancer: Advanced Character Animation from Video
Main Topics and Key Points:
- Steady Dancer is a new AI tool for generating dancing videos of any character.
- It takes a reference image of a character and a reference video of someone dancing to transfer movements.
- The model accurately transfers arm, body, and hand movements, even with different angles and character proportions.
- It works with both realistic and 3D fictional characters (e.g., Wukong).
- Comparison: Steady Dancer significantly outperforms previous leading tools like Juan Animate, offering more coherence, sharpness, and character consistency.
- Versatility: Beyond dancing, it can animate characters performing actions like sweeping leaves, shoveling snow, or kicking a soccer ball based on pose videos.
- Availability: The code is available on GitHub with instructions for local execution.
3. Viz Audio: Real-time Binaural Audio Generation from Video
Main Topics and Key Points:
- Viz Audio is an AI that generates binaural (3D) audio from silent videos.
- It understands the actions within a video to create spatially accurate soundscapes.
- Examples: Audio accurately follows the movement of objects (e.g., a tractor, a harpist, a trumpet) from right to left or across the screen. It can differentiate between instrument playing styles (slapping vs. strumming a guitar) and synchronize string sounds with bow movements.
- Architecture: Utilizes a dual branch audio generation technique, processing audio for left and right ears separately for realism. It also employs a conditional space-time module for action consistency.
- Availability: While no code is released yet, the developers plan to open-source the code, dataset, and training code.
4. AI Video Generation Advancements: Pixverse, Runway, Cling, and Hunyan
Main Topics and Key Points:
- Pixverse v5.5:
- Generates videos with native sound, up to 1080p resolution and 10 seconds duration.
- Feature: Multigeneration switch allows generating multiple videos of a scene from different angles for coherent micro-stories.
- Limitation: Dialogue can sound robotic and lack emotion.
- Runway Gen 4.5:
- Claims improvements in physics, motion, and prompt control.
- Limitation: Generations lack sound.
- Shows more coherence and camera control than previous versions but still has noticeable noise and artifacts, not on par with leading models like Vio or Sora.
- Cling (Two New Models):
- Cling01: An omnimodal model understanding text, images, and videos. It allows mixing and matching media for flexible editing (background/character replacement, object insertion). Described as "Nano Banana for video."
- Cling 2.6: The latest advanced model with native, good-quality, and in-sync sound. Excels in physics, high-action scenes, and image-to-video.
- Hunyan Video 1.5 (Tencent):
- An update to a popular open-source video model.
- Key Improvement: A new distilled model reduces generation time by 75%, allowing video creation in 8 or 12 steps instead of 50.
- Performance: On a single 4090 GPU, video generation takes approximately 75 seconds. The quality difference between the distilled and full model is negligible.
5. Engine AI: High-Speed Humanoid Robot Demonstrations
Main Topics and Key Points:
- Engine AI showcased impressive demos of their new T800 robot.
- The robot demonstrated high-speed spin kicks, punches, and multiple punch/kick combos, exceeding the speed and fluidity of robots like Optimus or Figure.
- Debate: It's unclear if the demonstrations were autonomous or teleoperated, but the speed and balance were notable.
- Verification: Behind-the-scenes footage was released to confirm the authenticity of the demos, dispelling claims of CGI.
- Potential Applications: While capable of combat-like movements, the robot's potential lies in industrial applications or search and rescue missions.
6. LongCat Image & Ovis Image: New Open-Source Image Models
Main Topics and Key Points:
- LongCat Image (Mateuan):
- A new open-source image generator with 6 billion parameters, similar to Z-Image.
- Showcases capabilities in creating posters, realistic photos, and various art styles.
- Testing: Initial tests revealed mixed results. While good at generating posters and some realistic images, it struggled with complex prompts involving text and specific scene elements (e.g., reflections, multiple objects).
- Image Editor: An accompanying edit model, similar to Nano Banana, was also released. Initial tests showed errors in editing floor plans and an inability to translate text.
- Strengths: Lightweight, open-source, and free to run.
- Ovis Image (Alibaba):
- A text-to-image model from a different Alibaba team, with 7 billion parameters.
- Strength: Excels at rendering text, outperforming other models in text generation benchmarks.
- Limitation: Image quality, especially for realistic photos, is not as high as Z-Image.
- Integration: Natively integrated into Comfy UI.
7. Flowith OS: Agentic Operating System for Autonomous Workflows
Main Topics and Key Points:
- Flowith OS is an agentic operating system where AI agents autonomously execute complex workflows across the web and within terminals.
- It operates across tabs, apps, and code environments, unlike simple AI chat features.
- Performance: Described as "shockingly smooth" and top-notch, outperforming other agentic browsers like Gemini Computer Use and ChatGPT Atlas.
- Capabilities: Automates tasks like liking and commenting on X (formerly Twitter) with thoughtful responses. Can handle command-line tasks, such as downloading and verifying a coding project from GitHub to a desktop without manual intervention.
- Availability: Currently in public beta, requiring an invite code.
8. Live Avatar: Real-time, Infinite-Length Video Generator
Main Topics and Key Points:
- Live Avatar (Alibaba) is a real-time video generator capable of producing infinite-length videos with audio.
- Technical Details: Achieves this through distribution matching distillation, transforming it into an efficient four-step streaming model, and timestep forcing pipeline parallelism (TPP) for parallel processing.
- Performance: Achieves an 84x FPS improvement over the baseline, generating videos at over 20 FPS.
- High-Fidelity Emotion Capture: Designed to mimic subtle emotions and nuanced expressions.
- Style Versatility: Generates realistic portrait videos (including full body movement), 3D Pixar-style, and 2D animations.
- Comparison: Significantly outperforms other video generators (Stable Avatar, Omni Avatar, Ditto) in maintaining quality over extended durations (minutes), where others show color changes, character distortions, or loss of functionality.
- Requirements: Currently requires five H100 GPUs for real-time generation.
- Future Plans: The team plans to release code, models, a Gradio demo, a model for low VRAM, and Comfy UI support.
9. Poster Copilot: AI-Powered Poster and Graphic Creation
Main Topics and Key Points:
- Poster Copilot is an AI agent that assists in creating professional posters and graphics.
- Input: Requires a prompt, dimensions (width/height), and optional assets/layers.
- Automation: Can automatically generate backgrounds or foreground elements if not provided.
- Iterative Refinement: Supports multi-round editing, allowing users to modify specific aspects of the poster (e.g., changing background color, material, camera angle) while maintaining text consistency.
- Aspect Ratio Conversion: Easily converts posters between different aspect ratios by reorganizing layers.
- Availability: Planned to be fully open-source, with code, data pipeline, dataset, and training code to be released.
10. Gemini 3 Deep Think: Advanced Reasoning Model
Main Topics and Key Points:
- Gemini 3 Deep Think is an enhanced version of Gemini 3, allocating more compute for extended reasoning.
- Capabilities: Excels at complex, multi-step questions in research, math, and coding.
- Benchmarks: Achieves top scores on the Humanity's Last Exam, graduate-level science and reasoning questions, and the Arc AGI 2 benchmark (visual puzzles testing pattern learning).
- Achievements: Awarded gold medal status in the International Math Olympiad and an international programming contest.
- Availability: Requires an Ultra subscription ($200+/month) and is not available for free or Pro users.
- Use Case: Best suited for professionals in fields requiring heavy thinking and research (medicine, math, programming). Considered overkill for everyday tasks.
11. Mistral 3: New Family of Open-Source Models
Main Topics and Key Points:
- Mistral AI released Mistral 3, a family of three small dense models (14B, 8B, 3B parameters) and a larger Mixture of Experts (MoE) model, Mistral Large 3 (675B parameters).
- Licensing: Released under the Apache 2.0 license, allowing commercial use.
- Performance: Self-reported benchmarks show Mistral Large 3 on par with older leading open-source models. Smaller models are competitive with similar-sized open-source alternatives.
- Independent Benchmarks: On the Artificial Analysis leaderboard (open-source only), Mistral Large 3 ranks lower than models like DeepSeek R1 or GLM 4.6.
- Recommendation: The smaller Mistral 3 models are good options for running locally on consumer-grade GPUs.
- Availability: Code and installation instructions are available.
12. DeepSeek V3.2: Top-Tier Open-Source Language Model
Main Topics and Key Points:
- DeepSeek V3.2 is presented as the best open-source model currently available, rivaling closed-source models like Gemini 3 Pro and GPT-5.
- Versions:
- Main v3.2: Successor to the experimental model.
- v3.2 Special: Optimized for maximum reasoning capacity.
- Performance: The special version rivals Gemini 3 Pro in reasoning tasks (math, science, coding). The regular v3.2 scores closely to proprietary models in agentic benchmarks.
- Achievements: The v3.2 Special achieved gold medal results in the International Math Olympiad (CMO), an international programming contest, and the International Olympiad on Informatics. This is significant as it's the first general open-source model to achieve this across all.
- Independent Ranking: On the Artificial Analysis leaderboard, it ranks second among open-source models, slightly behind Kim K2 Thinking.
- Cost-Effectiveness: Offers excellent value, with its native API costing approximately $0.30 per million output tokens, significantly cheaper than GPT-5.1 or Gemini 3 Pro.
- Availability: Live on app and web platforms. Hugging Face links are provided for download.
- Technical Details: A large 685 billion parameter model, approximately 690 GB in size, requiring enterprise-grade hardware for local deployment.
13. Cadream 4.5: State-of-the-Art Image Generation and Editing
Main Topics and Key Points:
- Cadream 4.5 is a new state-of-the-art image model from ByteDance, improving upon previous versions.
- Text-to-Image: Generates highly detailed and realistic photos, excelling in aesthetics. It handles complex scenes, dynamic motion (long exposure), and various art styles (e.g., watercolor).
- Text, Posters, Graphics: Exceptionally good at generating text, posters, and graphics, often creating multiple well-designed posters in a single generation. It can also generate entire user interfaces with correct text.
- Image Editing: Functions as an image editor similar to Nano Banana. Can remove characters, change image themes (e.g., to nighttime), translate text (e.g., to Chinese in a handwritten font), change text attributes (color, italics), and integrate images into posters.
- Limitations: Closed-source and proprietary, not available for free offline use.
14. Tuna: Unified Multimodal Model from Meta
Main Topics and Key Points:
- Tuna is a unified multimodal model from Meta that understands and generates text, images, and video.
- Capabilities:
- Video Generation: Produces videos (384x672 resolution, 12 FPS), described as a proof of concept rather than state-of-the-art.
- Image Generation: Creates realistic images, including different camera perspectives and text within images.
- Image Editing: Performs image editing tasks, with examples showing it outperforming open-source competitors like Bagel, Quen, and Flux One in specific edits (e.g., changing doll orientation, applying lighting, style transfer, object replacement).
- Chatbot Functionality: Acts as a chatbot that can analyze uploaded images and videos, answering questions accurately.
- Availability: Code is under legal review and not yet released. Given Meta's history, public release is uncertain.
15. Lotus 2: Advanced Depth and Normal Estimation
Main Topics and Key Points:
- Lotus 2 is an AI model capable of predicting image depth and normal (surface orientation).
- Performance: Significantly outperforms competitor models in depth estimation, capturing fine details in both foreground and background. It also provides more detailed normal estimations, including subtle surface details like fabric.
- Architecture: Uses a two-stage process: a core predictor for rough 3D shape estimation, followed by a detail sharpener to refine predictions.
- Availability: Code is available on GitHub with instructions for local execution.
- Requirements: Requires at least 40 GB of VRAM and uses Flux One, which is noted as an older technology compared to current open-source alternatives.
Conclusion and Key Takeaways
The past week has been exceptionally productive in the AI landscape, marked by a rapid succession of releases across various domains. The trend towards more accessible, real-time, and multimodal AI continues, with significant strides in open-source models challenging the dominance of proprietary ones. Key takeaways include:
- Real-time capabilities are becoming standard for TTS and video generation, enhancing user experience and interactivity.
- Open-source models like DeepSeek V3.2 are achieving performance levels that rival or surpass leading closed-source alternatives, democratizing access to advanced AI.
- Multimodality is a growing focus, with models like Cling01 and Tuna demonstrating the ability to process and generate multiple data types within a single framework.
- Efficiency and accessibility are being addressed through techniques like model distillation and the development of smaller, consumer-grade compatible models.
- Specialized models for complex reasoning (Gemini 3 Deep Think) and specific tasks (Poster Copilot, Viz Audio) are emerging to cater to niche but important applications.
- The humanoid robot sector is showing remarkable progress in speed and agility, hinting at future real-world applications.
The rapid pace of innovation suggests that the AI field will continue to evolve at an accelerated rate, with further breakthroughs expected in the coming weeks and months.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

What's new in Angular
Chrome for Developers

Rachel Reeves tells Sky News she will still be chancellor for the autumn budget
Sky News

Inside the Enhanced Games, aka 'The Doping Olympics' | The Global Story
BBC News

'DHS OFFICIAL ORDERED ME TO DELETE…': Witness reveals SHOCKING details of Minnesota Child Care fraud
The Economic Times

Penélope Cruz and Glenn Close star in Spanish civil war gay drama at Cannes • FRANCE 24 English
FRANCE 24 English

Getting More from Every Copilot Interaction
GitHub

U-Haul trucks are turning around. The Exodus is OVER.
Reventure Consulting