New robot waifus, GLM 5.2 craze, AI spas, new world models, new science agents: AI NEWS
By AI Search
Share:
Key Concepts
- World Models: AI systems (e.g., DreamXWorld) that generate explorable, interactive 3D environments from prompts.
- Persistent Memory: A mechanism in video AI (e.g., Permavid) to maintain object and structural consistency during edits.
- Unified Scientific Models: AI (e.g., Logos) that process diverse scientific data (proteins, molecules) using a shared grammar.
- Exoskeleton Robot Control: Wearable systems (e.g., Universal Manipulation Exoskeleton) for teaching robots via human movement and force feedback.
- Model Compression: Techniques (e.g., GGUF, 1-bit/2-bit quantization) to run massive models on consumer hardware.
- Agentic Workflows: AI systems (e.g., OpenAI’s Record and Replay) that learn complex tasks by observing human screen recordings.
1. Generative Video and World Models
- DreamXWorld: A new world model trained on Unreal Engine data and real-world video. It allows users to prompt an environment that evolves based on actions (e.g., driving, riding). It supports long-form video generation with high temporal consistency.
- Permavid: Addresses the "consistency problem" in video editing. It utilizes two memory banks—one for appearance and one for 3D structure—allowing users to perform global style changes or local object edits without losing scene geometry.
- Omni Director (Kling): A system that clones camera motion from a reference video onto a new source image. It supports complex movements like dolly zooms, bullet time, and multi-shot sequences.
2. Image Generation and Editing
- Boo Goo Image: An open-source image generator and editor under the Apache 2 license. It excels at text rendering and infographic generation. While it offers a "Turbo" mode for speed, it is noted to be slower and less photorealistic than Flux or Z-Image in some benchmarks.
- Telestyle V2: A style transfer tool capable of handling diverse combinations of content and style images, overcoming the limitations of older models that required specific "realistic-to-artistic" pairings.
3. Robotics and Physical AI
- Universal Manipulation Exoskeleton (Ant Group): A wearable system that records human arm motion and force/torque feedback. This data is used to train robots to perform household tasks (e.g., opening a fridge) with physical awareness.
- Table Tennis Robots:
- Sony’s "Ace": A high-speed robot mounted on a motorized rail that uses real-time spin detection to compete against professional human players.
- AGI Bot A3: A humanoid robot utilizing the "Spike Ping Pong" algorithm, enabling 10x faster vision response and millimeter-level precision while maintaining bipedal balance.
- Moya (Droid Up): A full-body humanoid robot designed for companionship and elderly care, capable of basic household chores.
4. AI for Science
- Logos (Alibaba): A unified model for scientific discovery. By tokenizing data from proteins, small molecules, and chemical reactions, it provides a shared framework for material generation and antibody design.
- AI Chemist (OpenAI/Maria): An autonomous system that successfully identified "TEMPO" as an additive to improve the yield of the Chan-Lam reaction, demonstrating AI's ability to navigate the scientific research loop (propose, test, analyze, refine).
5. Large Language Models (LLMs)
- GLM 5.2: Currently the leading open-source model. It ranks highly on the "Artificial Analysis" intelligence index, outperforming many frontier models in factual accuracy with a significantly lower hallucination rate.
- Accessibility: Despite the full model being 1.5 TB, the community (via Unsloth) has released compressed GGUF versions (1-bit and 2-bit), allowing the model to run on high-end consumer hardware (e.g., RTX 6000s or Mac Studio).
6. Automation and Workflow
- Record and Replay (OpenAI): A feature for the Codex agent that learns skills by watching screen recordings of a user performing a task. It is designed for stable, repetitive workflows like metadata entry or file management.
7. Midjourney Pivot: Midjourney Medical
- Concept: Midjourney is moving into health technology with a "spa-based" body scanning system.
- Technology: Uses underwater ultrasound sensors (half a million tiny elements) to create 3D maps of the body in 60 seconds.
- Timeline: A research spa is planned for San Francisco in 2027, with a focus on body composition tracking before seeking FDA approval for broader medical diagnostics.
Synthesis
The current AI landscape is shifting from simple generation to persistent, agentic, and physical integration. The emergence of models like GLM 5.2 and tools like Permavid demonstrates that the industry is solving the "consistency" and "hallucination" bottlenecks. Simultaneously, the integration of AI into physical robotics (Sony Ace, Ant Group Exoskeleton) and scientific discovery (Logos, AI Chemist) signals a transition toward AI that can interact with and manipulate the real world, rather than just generating digital content.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

Nvidia Wants to Make Humanoid AI Robots Safer Around Humans
Bloomberg Technology

China's New AI Robot MOYA Feels Too Real (92% Human)
AI Revolution

Claude Code WordPress SEO: Automate Everything ($500K+ Earned)
Jono Catliff

The Best AI Automation Stack to Learn in 2026
Dave Ebbelaar

India goes football crazy: Is politics holding the country back? • FRANCE 24 English
FRANCE 24 English

Can This $125K Robot Be My Friend? | Big Business
Business Insider

Mỹ tiêu diệt trùm băng đảng khét tiếng tại Venezuela | Cụm tin | VTV24
VTV24