Realtime AI video, open-source SUNO, next-level AI agents, realtime text-to-speech: AI NEWS
By AI Search
AI Weekly Update: Recent Developments & Tools
Key Concepts: AI Agents, Real-time Video Generation, 3D Modeling, Audio Enhancement, Image Generation/Editing, Text-to-Speech, Open-Source Models, LLM (Large Language Model), VRAM (Video Random Access Memory), SLAM (Simultaneous Localization and Mapping).
I. AI Agents & Automation – Show UI Aloha & Show UI Pi
This week saw significant advancements in AI agents capable of interacting with computer interfaces. Show UI Aloha is a novel AI agent that learns by observing human task completion. It records user actions (clicks, drags, typing, scrolling) and translates them into code, enabling autonomous task replication. Examples include booking flights, transposing matrices in Excel, and batch editing in PowerPoint. The architecture involves a Recorder component, a Learner, a Teach Trajectory generator, and an Executor. Performance benchmarks demonstrate superiority over existing agents like UI, TARS, and Claude for Sonnet. The code is available on GitHub (link in description), requiring an LLM provider API key for operation.
A complementary agent, Show UI Pi, focuses on smooth mouse movement control. Unlike other agents struggling with drag-and-drop actions, Show UI Pi excels at tasks like solving slider captchas, drawing, and manipulating video editor clips. It even demonstrates proficiency in handwriting and solving complex captchas, outperforming Gemini in PowerPoint and video editing tasks. While documentation is currently incomplete, the code is available on GitHub with intentions for full open-sourcing.
II. Audio Processing – Nova SR
Nova SR is a remarkably small (50KB) AI model for real-time audio upscaling. It achieves over 3,500x real-time speed on standard hardware, significantly improving audio quality. Demonstrations showcase enhanced clarity in spoken word and music. The model’s performance is comparable to models 5,000 times larger, running efficiently on consumer GPUs. Nova SR is open-source, with code available on GitHub and a free Hugging Face Space for online testing.
III. Video Generation & Manipulation – Verse Crafter & Pixver R1
Verse Crafter (Tencent) generates videos from single images, offering precise control over camera movement and object manipulation. It utilizes MO and SAM 2 to create a 3D point cloud from the input image, which can then be edited in Blender. Users can define camera trajectories and object movements, resulting in highly customizable videos. Verse Crafter outperforms competitors like Ume and Uni3C in coherence and quality. The GitHub repository is publicly available, including a Blender add-on.
Pixver R1 is a real-time world model capable of generating interactive scenes from text prompts. While currently exhibiting inconsistencies and visual glitches (demonstrated with a woman walking in a city and subsequent explosions), it represents a significant step towards continuous, prompt-driven video generation. It allows for dynamic scene modification with ultra-low latency.
IV. 3D Modeling – Shape R & UniSH
Shape R (Meta) creates metric-accurate 3D models of scenes from videos or images, generating individual 3D models for each object. This contrasts with traditional methods that produce a single, monolithic model. It leverages SLAM and multimodal conditioning for accurate reconstruction, even from regular phone footage. Code and data are available on GitHub.
UniSH is an AI designed to automate the rigging process for 3D models, adding a skeleton to enable animation. It can handle various subjects, including humans, animals, and inanimate objects, and outperforms existing methods like UniG. Code and data are forthcoming.
V. Image Generation & Editing – Flux 2 Klein & Vibe
Flux 2 Klein is a fast image generator and editor, comparable to Nano Banana. While image generation quality is close to Zimage, Flux 2 Klein excels in image editing, often surpassing Quen Imageit 2511 in speed. However, it struggles with anatomical accuracy.
Vibe is another blazing-fast open-source image editor capable of generating 2K images in just 4 seconds (requiring 24GB VRAM). It combines Nvidia’s Sauna 1.5 and Quen 3VL for efficient editing, allowing for tasks like adding objects, changing backgrounds, and applying artistic styles. It’s available on GitHub and Hugging Face with a demo.
VI. Image Depth Estimation – AnyDepth
AnyDepth is a new depth estimator that enhances the detail of depth maps generated by models like Depth Anything. It achieves higher resolution and fidelity with lower latency than competitors like Dino and DPT. It’s based on a Simple Depth Transformer architecture and is available on GitHub with a prepared dataset.
VII. Text-to-Speech – TinyTTS
A new tiny text-to-speech generator (100 million parameters) can run in real-time on a CPU. It offers ultra-low latency (200ms) and voice cloning capabilities, currently supporting only English. Code is available on GitHub.
VIII. Image to Video – AI-Driven Movement Generation
An unnamed AI tool demonstrated the ability to generate videos from image movements. By dragging elements within an image, the AI creates a video depicting those movements.
IX. LoveArt – AI Design Platform (Sponsored)
LoveArt is an AI-powered design platform offering features like flash sale poster generation, object manipulation, and background editing. Its "Edit Elements" feature allows for segmented editing of images, and its "Edit Text" feature enables precise text modifications. It also supports video generation and is positioned as a comprehensive AI design solution.
Conclusion:
This week’s AI developments showcase rapid progress across multiple domains. The trend towards smaller, faster, and more accessible open-source models is particularly notable, exemplified by Nova SR, Vibe, and TinyTTS. Advancements in AI agents (Show UI Aloha & Pi) and video generation (Verse Crafter, Pixver R1) are pushing the boundaries of automation and creative content creation. The release of tools like Shape R and UniSH further democratize 3D modeling and animation. The continued evolution of these technologies promises to significantly impact various industries and creative workflows.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

Stanford CS153 Frontier Systems | Building the Frontier Ecosystem
Stanford Online

'Things are going to be okay, in Canada and the U.S.': Thorne
BNN Bloomberg

I'M OUT: The $11 Trillion AI Bubble is Breaking!
Steven Van Metre

South Korea bets big on AI with nearly a trillion dollars of investment • FRANCE 24 English
FRANCE 24 English

The Bubble is Bursting... (Emergency Update)
Bravos Research

The AI Bubble Just Ended - Without Popping
Heresy Financial

AI Market Volatility, Europe Heat Wave, Venezuela Quakes Damage | Bloomberg This Weekend: June 27
Bloomberg Television