Open-source Sora 2, realtime AI video, new DeepSeek, Nanobanana upgrades, Claude 4.5, realtime TTS
By AI Search
Key Concepts
- One Alpha (Alibaba): Video generation with native transparency.
- Canyt: Real-time, low-latency, open-source text-to-speech.
- Cap 4D: Real-time 4D (3D + time) head avatar creation.
- Gemini 2.5 Flash (Nano Banana): Image generation with new aspect ratio control.
- Hunyan Image 3.0 (Tencent): Open-source image generator with "world understanding," strong text generation, and aesthetics.
- Long Live (Nvidia): Real-time interactive video generation with prompt-based editing.
- OVI (Character AI): First open-source video generator with native audio, supporting text-to-video and image-to-video.
- Omni Retarget: Robot learning from human motion capture data for complex tasks.
- GLM 4.6 (ZAI): Latest open-source large language model (LLM) with increased context, superior coding, and agentic capabilities.
- Light X2V 2.2 Lightning LoRA: Accelerator plugin for local AI video generation (e.g., One 2.2).
- Claude Sonnet 4.5 (Anthropic): Latest LLM, claimed "best coding model," but shows mixed performance in practical tests.
- Deepseek 3.2 Experimental: Efficient open-source LLM, optimized for efficiency over raw performance.
- Dreamer 4: AI agent that trains itself in a simulated environment (e.g., Minecraft) to learn complex tasks.
- Transparency Generation: Ability to create video content with an alpha channel for seamless layering.
- Streaming Long Tuning: Technique for efficient and consistent long video generation.
- Mixture of Experts (MoE): LLM architecture where multiple "expert" models contribute to a response.
- Quantization/Compression: Reducing model size and computational requirements for consumer-grade hardware.
- LoRA (Low-Rank Adaptation): A method to fine-tune large models efficiently, often used as plugins.
- SWEBench Verified Benchmark: A specific benchmark for evaluating coding capabilities of LLMs.
- DeepSeek Sparse Attention: Mechanism to improve training and inference efficiency in LLMs.
- World Model/Simulation: An AI's internal representation of an environment used for self-training.
Alibaba's One Alpha: Video Generation with Transparency
Alibaba has released One Alpha, a new model that generates videos with native transparency. This builds on their previous release, Juan 2.5, which was a VO3/Sora 2 competitor capable of generating video with audio. One Alpha's key feature is its ability to produce videos where the background is transparent, allowing users to easily layer the generated content onto any existing background.
The model demonstrates impressive capability in handling various levels of transparency, from translucent objects like bubbles and glass bottles to complex elements such as leaf foliage and hair, which are notoriously difficult to segment accurately. For effects like glowing flames or halos, One Alpha can generate gradual transparency, as evidenced by its accurate alpha channel output. All models are available on HuggingFace, and a GitHub repository provides instructions for local download and execution.
Canyt: Real-time Text-to-Speech
Canyt is a new open-source text-to-speech (TTS) model that boasts real-time generation capabilities. It can generate 15 seconds of audio in just 1 second using an RTX 5080 GPU, requiring only about 2 GB of VRAM, which translates to super low latency.
The model is remarkably tiny, with only 370 million parameters, significantly smaller than some TTS generators that exceed a billion parameters. It supports multiple languages and operates under the Apache 2 license, allowing for commercial use with minimal restrictions. Canyt's architecture involves using its own language model to convert transcripts into audio tokens, which are then transformed into smooth, natural-sounding audio by Nvidia's Nano Codec.
A free HuggingFace Space is available for online trials, offering multiple predefined speakers and language options (e.g., English, Irish, Arabic, Korean). Demos show its ability to understand context, pronounce tricky words, and even tongue twisters correctly. The GitHub repository provides instructions for local download and execution.
Cap 4D: Real-time 3D Head Avatars
Cap 4D is an AI tool that creates real-time 4D avatars, meaning 3D faces that can move and change over time (the fourth dimension). Unlike previous 2D avatar generators, Cap 4D constructs a real-time 3D model that allows for head rotation and control over facial expressions. Users can upload multiple images of a person from different angles to construct the avatar, with more images leading to higher accuracy. The code for Cap 4D has recently been released on a GitHub repository, including instructions for downloading and operating these 4D avatars locally.
Gemini 2.5 Flash (Nano Banana) Update: Aspect Ratio Control
Google's Gemini 2.5 Flash, also known as Nano Banana, has received an update that allows users to control the aspect ratio of output images. Previously, this feature was unavailable. The model now offers 10 predefined aspect ratios. This functionality is accessible for free within Google's AI Studio, where users can select Nano Banana and choose their desired aspect ratio before generating images. An example demonstrated changing a horizontal input image to a vertical output image while applying a prompt.
Tencent Hunyan Image 3.0: Top Open-Source Image Generator
Tencent Hunyen has released Hunyan Image 3.0, a powerful open-source image generator with advanced "world understanding," akin to GPT's image generator, Google's Imagine, and Quen Image.
Key Strengths:
- World Understanding: Capable of generating complex scenes, multi-panel tutorials (e.g., a 9-square grid on sketching a parrot), and intricate illustrations (e.g., evolution from monkeys to humans at a computer).
- Aesthetics: Produces realistic and "unrefined" images, avoiding the "plasticky" look of some earlier models like Flux.
- Text Generation: Excels at generating text within images, including complicated multi-material letters and long snippets, making it suitable for infographics, diagrams, and presentation slides.
Benchmarks and Comparisons: Hunyan Image 3.0 was benchmarked against closed-source models like Cream, Nano Banana, and GPT Image. While it showed comparable overall performance to these top models with English prompts, it demonstrated significantly better performance in some categories when using Chinese prompts. Professional evaluators also gave it the highest win rate among the compared models.
User Testing and Limitations:
- Pokemon Card Generation: Successfully generated most specified details (HP, moves, damage) but produced gibberish for unspecified text.
- Anime Characters: Accurately depicted popular anime characters (Naruto, Nezuko, Goku, Doraemon, Emilia, Satoru Gojo) taking a selfie, including accurate logos, though some outfit details were incorrect.
- Data Visualization: While it correctly generated text for a presentation slide (title, subtitle, source, legend), it struggled to create an accurate pie chart from provided data, showing incorrect order and extra labels. It appears to handle simple presentation slides but fails with complex data visualization.
- Meme Grid: Generated a 3x3 grid of student life memes with specified themes, but had subtle flaws like gibberish text and distorted eyes in some panels.
- Spot the Difference: Successfully created two side-by-side panels with subtle differences.
- Shadow Puppet: Failed to realistically generate a hand making a shadow that looked like a rabbit.
Availability and Hardware: Hunyan Image 3.0 can be tried for free via its online interface. While it is open-sourced for local download, it has extremely high hardware requirements: 170 GB of storage and 240 GB of VRAM (e.g., three 80 GB GPUs), making it impractical for most consumer-grade setups. Tencent is actively working on a quantized (compressed) version to run on more accessible hardware.
Nvidia Long Live: Real-time Interactive Video Generation
Long Live by Nvidia is an AI that generates real-time interactive videos. Users can input a prompt to generate a video and then inject further prompts to edit the video as it plays, with very low latency for applying changes.
While the video quality is not as high as leading non-real-time models (exhibiting noise, warped faces, and fingers), this is expected for a real-time generator. A significant achievement is its ability to generate videos up to 4 minutes long.
Technical Details: Long Live employs a technique called "streaming long tuning," which breaks the video generation process into smaller chunks and reuses previously generated frames to enhance efficiency and consistency. The model is based on One 2.1 (1.3 billion parameters). A GitHub repository is available with instructions for local download and execution. It was tested with an Nvidia GPU with 40 GB of memory and 64 GB of RAM, but other setups may work. It supports approximately 21 frames per second (FPS) generation on a single H100 GPU and 24.8 FPS with a quantized version, with marginal quality loss.
OVI by Character AI: Open-Source Video with Native Audio
OVI by Character AI is a groundbreaking open-source video generator that includes audio natively built in. This distinguishes it from closed-source competitors like Sora 2, V3, and Juan 2.5. OVI supports both text-only prompts and image-to-video generation, allowing users to upload a reference image as the starting frame, even for realistic human subjects—a capability Sora 2 lacks for such scenarios.
Key Features:
- Audio Control: Prompts can specify dialogue, desired voice characteristics, and background noise.
- Multiple Speakers: Capable of handling multiple speakers within a scene, a challenge for some other models.
- Multilingual and Expressive: Works with different languages and expressions.
- Singing Capabilities: Appears to have built-in music generation capabilities.
Availability and Hardware: OVI is already released, with a GitHub repository providing instructions for local download and execution. The minimum GPU requirement is 32 GB, and the open-source community is expected to further quantize the model for lower VRAM systems.
Omni Retarget: Robot Motion Learning from Humans
Omni Retarget is an impressive AI that teaches robots how to perform complex movements and tasks by copying human motion capture data. The process involves using human motion data (e.g., jumping, moving boxes, flips, parkour) as input, mapping these movements onto a robot, and then fine-tuning the robot's actions with reinforcement learning to ensure balance and stability.
The results are remarkable, with robots performing acrobatic feats like jumping off tables, rolling, climbing, mid-air rotations, crawling, and carrying boxes with natural, human-like movements. These actions are performed autonomously, not via tele-operation. The system can learn complex sequences of tasks up to 30 seconds long, such as a robot moving a chair, stepping on it to climb a table, jumping down, and performing a flip. A dataset for this technology has been released, and the code is planned to be open-sourced soon.
ZAI GLM 4.6: Latest Open-Source LLM
ZAI has released GLM 4.6, their latest open-source large language model, building on the strong agentic performance and coding capabilities of previous GLM models.
Notable Updates in 4.6:
- Context Window: Increased from 128K to 200K tokens, allowing for more information in a single prompt.
- Coding Performance: Achieves superior scores on coding benchmarks.
- Reasoning and Agentic Capabilities: Enhanced advanced reasoning and agentic performance.
- Writing Quality: Aligns better with human preferences in style and readability, and performs more naturally in role-playing scenarios.
Benchmarks and Efficiency: GLM 4.6 performed exceptionally well across various benchmarks, including competitive math (AIM), graduate-level science questions (GPQA), humanities exams, coding, and agentic tasks, often outperforming Deepseek 3.2 experimental and Claude Sonnet 4.5. It also boasts high efficiency, using the least amount of tokens when executing tasks, resulting in a higher win rate compared to other models.
User Demos:
- Color Palette Generator: Successfully generated functional HTML code for a tool that extracts colors from uploaded images, suggests harmonies (complementary, analogous, triadic), and allows palette export.
- CRM Dashboard: Created a beautiful, real-time updating HTML CRM dashboard with interactive charts (funnel, pie, heat map), sales rep performance, and activity logs. The dashboard elements were draggable and repositionable.
- E-commerce Growth Analysis (with web search): Provided a comprehensive executive summary, key metrics, transaction trends, market penetration analysis, regional analysis, growth drivers, and future outlook. While not as detailed as dedicated deep research agents, it performed well for a base LLM, though it failed to generate a map with highlighted countries.
Availability and Hardware: GLM 4.6 is available for free via Z.AI's chat interface. It is also open-sourced for local download as a 335 billion parameter Mixture of Experts (MoE) model, with only 32 billion parameters active during use for efficiency. However, it requires substantial hardware (e.g., 8 H100s or 4 H200s for the FP8 version of GLM 4.5), making local execution challenging for most users without GPU cloud rental.
Light X2V 2.2 Lightning LoRA: Video Generation Accelerator
For users generating AI videos locally with models like One 2.2 or 2.1, the Light X2V plugin acts as a LoRA (Low-Rank Adaptation) to accelerate video generation by up to 20 times. It reduces the required step count from 20-30 to just 4 steps.
A new version, Light X2V 2.2 Lightning LoRA, has been released, offering significant improvements over its predecessor (version 1.1). The latest version produces much more realistic and natural movements, faster actions, and more accurate physics, addressing issues like slow-motion generation and incorrect physics seen in earlier versions. These Loras are available for download from a HuggingFace folder and can be integrated into One 2.2 workflows.
Anthropic Claude Sonnet 4.5: Latest LLM
Anthropic has released Claude Sonnet 4.5, their latest and most advanced model, aggressively claiming it to be the "best coding model in the world." They assert substantial gains in reasoning, math, coding, and agentic capabilities, noting its ability to focus for over 30 hours on complex multi-step tasks.
Benchmarks and Caveats: Claude 4.5 leads the SWEBench verified benchmark with an 82% score, surpassing GPT5 and GPT5 codecs. However, this score is achieved using "parallel test time compute" (generating multiple attempts and selecting the best), a method not applied to competitor models in the comparison, making the claim potentially misleading. On other benchmarks like Terminal Bench Hard, science-related tasks (Humanity's Last Exam, GPQA Diamond), live codebench instruction following, and competitive math, Claude 4.5 Sonnet performs less favorably, often falling behind competitors like Grok 4, GPT5 codecs, and even older models like Gemini 2.5 Pro.
User Demos and Practical Performance: User tests using complex coding prompts revealed limitations:
- Real-time Ray Tracing Simulation (Metallic Sphere): Claude 4.5 generated a basic interactive scene with adjustable parameters, but failed to import a real 3D street view map and its reflectivity parameter seemed ineffective, performing worse than GPT5.
- Beehive Construction Simulation: The generated visualization was inaccurate, with improperly positioned hexagonal cells, no expansion, and bees not foraging, again falling short compared to GPT5.
- Fluid Dynamics Animation: The animation did not accurately represent fluid mixing, and interactive sliders failed to produce realistic fluid behavior, contrasting with GPT5's superior execution.
Conclusion: Based on practical tests, the user was not impressed by Claude Sonnet 4.5's coding capabilities, suggesting it struggles with many real-world coding tasks despite its benchmark claims. For closed-source coding models, the user still prefers Grok 4 or GPT5.
Deepseek 3.2 Experimental: Efficient Open-Source LLM
Deepseek has released Deepseek 3.2 Experimental, an experimental version building upon their leading open-source model, 3.1 Terminus. This new version refines the model by introducing the "DeepSeek sparse attention" mechanism, which significantly improves training and inference efficiency, particularly for long context scenarios, leading to a substantial reduction in cost compared to 3.1 Terminus.
Performance vs. Efficiency: Deepseek 3.2 is primarily optimized for efficiency rather than raw performance. While it performs better in some benchmarks (competitive math, code forces, agentic tasks), it may not perform as well in others (e.g., scientific knowledge) compared to 3.1 Terminus. Overall, its performance is considered on par with 3.1 Terminus, but with much greater efficiency. On an independent leaderboard, 3.1 Terminus tied for first place, while 3.2 tied for second, but at a significantly lower cost. The model is available on HuggingFace with instructions for local download.
Dreamer 4: Minecraft AI Agent Training in Simulation
Dreamer 4 is a highly complex AI agent that trains itself to play Minecraft by "imagining" gameplay within its own simulation, rather than through real-world trial and error. The process involves:
- Observation: Watching extensive Minecraft gameplay videos.
- World Model Construction: Reconstructing its own internal "world model" or simulation of Minecraft.
- Self-Practice: Practicing gameplay within this simulation, trying thousands of actions to learn what works (e.g., repeatedly chopping trees to gather wood, mining stones).
Key Achievement: Dreamer 4 is the first AI agent to successfully mine diamonds in Minecraft, a notoriously difficult task requiring over 20,000 mouse and keyboard actions. It first rehearsed this complex sequence in its simulation before successfully achieving it in an actual Minecraft game, albeit with a 7% success rate.
Significance and Real-World Applications: While Minecraft is the initial application, the underlying principle of Dreamer 4 is profound. It suggests a future where AI models can train robots for highly complex, multi-step tasks by having the AI imagine and simulate the task internally, thereby reducing the time and resources required for physical training in the real world. Currently, only a technical paper has been released, with no public code available.
Synthesis and Conclusion
This week in AI has been marked by an extraordinary pace of innovation across diverse domains. We've seen significant advancements in video generation, with Alibaba's One Alpha introducing native transparency for seamless layering and Character AI's OVI pioneering the first open-source video model with native audio, even handling image-to-video for realistic human subjects. Real-time capabilities were a recurring theme, exemplified by Canyt's efficient text-to-speech and Nvidia's Long Live enabling interactive video editing on the fly.
In image generation, Tencent's Hunyan Image 3.0 emerged as a powerful open-source contender, showcasing impressive "world understanding" and text generation, though its local deployment remains challenging due to extreme hardware demands. Large Language Models (LLMs) continued their rapid evolution, with ZAI's GLM 4.6 demonstrating superior coding and agentic performance with an expanded context window, and Deepseek 3.2 focusing on efficiency improvements. Anthropic's Claude Sonnet 4.5, despite aggressive marketing for coding, showed mixed practical performance compared to established competitors.
Beyond content creation, AI for robotics made strides with Omni Retarget enabling robots to learn complex human movements from motion capture data. Perhaps most forward-looking is Dreamer 4, an AI agent that trains itself for highly complex tasks like diamond mining in Minecraft through internal simulation, hinting at a future where robots could learn intricate real-world tasks without extensive physical training. The continuous release of open-source models and the focus on real-time, efficient, and autonomous AI systems underscore a dynamic and competitive landscape, pushing the boundaries of what AI can achieve.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

How the hometown humiliation of Putin marks a turning point for Ukraine | DW News
DW News

Shocking video shows moment paramedics are hit by Israel in 'double-tap' strike
Sky News

Good Morning America Full Broadcast - Saturday, May 23, 2026
ABC News

Putin Xi, To Catch a Castro, Red Carpet Rebellion • FRANCE 24 English
FRANCE 24 English

Is the World Cup becoming the Superbowl?
CGTN America

New York Knicks super-fan Spike Lee talks movies and his love of basketball
ABC News

Kyle Busch Dies; Travel Restrictions Over Ebola Outbreak. What You Need to Know - May 22, 2026
ABC News