Opus 4.6, GPT 5.3 Codex, StepFun, Qwen3 Coder, new deepfake AIs, new video tools: AI NEWS

By AI Search

Share:

Key Concepts

  • Rapid advancements across multiple AI domains including OCR, avatar creation, LLMs, video editing, and robotics.
  • Increasing competition and accessibility with both powerful closed-source and competitive open-source AI models.
  • Emergence of AI tools capable of sophisticated video editing, particularly in speech manipulation, raising ethical concerns about deepfakes.
  • Recursive self-improvement demonstrated by new AI models, like GPT 5.3 Codeex, contributing to accelerated development.

AI Developments - This Week’s Highlights

This week witnessed a surge of development in Artificial Intelligence, spanning Optical Character Recognition (OCR), interactive avatars, Large Language Models (LLMs), video editing, and robotics. A significant trend is the release of both powerful closed-source models and competitive open-source alternatives.

Optical Character Recognition (OCR)

ZAI released GLM OCR, a new OCR model that outperforms existing open-source (BU’s Paddle Plus, DeepSc OCR 2) and closed-source (Gemini 3 Pro, GPT 5.2) models in parsing text, tables, formulas, and handwritten data from images. GLM OCR, at 2.6GB, enables local execution on consumer-grade hardware and is available on Hugging Face. It demonstrates faster processing speeds and higher accuracy across benchmarks including code parsing, table extraction, handwriting recognition, and multi-language support.

Interactive Avatars

Tencent’s Interact Avatar allows for the creation of talking characters capable of interacting with objects based on text prompts – examples include picking up and using headphones, phones, cameras, and plush toys. Utilizing audio input (though quality impacts realism), the system generates complex actions and gestures, even specifying precise timing. It surpasses existing animation competitors like Omni Avatar by enabling object manipulation. The code is available on GitHub, utilizing 12.2 as the base video model.

Large Language Models (LLMs)

Several notable LLMs were released or updated:

  • Anthropic’s Claude Opus 4.6: Anthropic’s most powerful model to date, outperforming GPT 5.2 in knowledge work, agentic search, and reasoning with tools. Surprisingly, it scored lower on the SUI Bench Verified benchmark compared to Opus 4.5, but significantly higher on the Arc AGI 2 benchmark (68.8% vs. GPT 5.2), which tests the ability to solve novel visual puzzles. Independent leaderboards (LM Arena, artificial analysis) also rank Opus 4.6 as top-performing, though it is slower and more expensive than competitors.
  • OpenAI’s GPT 5.3 Codeex: Released shortly after Claude Opus 4.6, GPT 5.3 Codeex is a highly capable coding agent. OpenAI claims it assisted in its own development, demonstrating recursive self-improvement – “A GPT 5.3 Codeex is the first model that was instrumental in creating itself.” It outperforms previous Codeex versions on SWE Pro and Terminal Bench, and is reported to be extremely powerful for autonomous coding tasks, even generating playable 3D games and Microsoft Office documents. It surpasses Opus 4.5 on SweetBench Pro and significantly outperforms it on Terminal Bench 2.
  • Step 3.5 Flash: A 200 billion parameter (11 billion active) open-source model rivaling top closed-source models in intelligence. It excels in deep reasoning and boasts a generation throughput of 100-300 tokens per second. It performs comparably to Gemini 3 Pro and GPT 5.2 on reasoning and coding benchmarks, and even surpasses them in scientific research. The model size is 399GB, requiring substantial hardware.
  • MiniCPM04.5: An open-source, omnimodal model (text, image, audio, video) capable of real-time interaction and voice conversion. It can parse text from images, answer reasoning questions based on visual input, and perform tasks like image recognition. Its size is 23.4GB, making it accessible on consumer hardware.
  • Quen 3 Coder Next: An 80 billion parameter open-source coding agent from Alibaba, demonstrating performance comparable to larger models like DeepSeek and GLM 4.7 on coding benchmarks. It’s particularly efficient and can be integrated with tools like Cloud Code and Open Claw.

Video Editing Advancements

Significant progress was made in video editing AI:

  • Omnimat Zero: This AI can remove objects from videos, including reflections, and seamlessly replace backgrounds, outperforming existing object removal tools. The code is available on GitHub.
  • Edit Yourself: This AI allows for editing video dialogue, adding or removing segments, and automatically adjusting lip sync to match the new audio. This raises concerns about potential misuse for creating deepfakes. It is available on foul.ai and Pipio.
  • Fast VMT: This AI transfers motion from one video to another, allowing for scene changes while preserving movement, three times faster than existing methods.
  • Context Forcing: A framework for generating longer, more consistent videos, extending generation lengths by 2-10x with minimal errors.

Robotics

  • Husky (Uni G1): Researchers developed a framework enabling the Uni G1 robot to ride a skateboard, demonstrating advanced balance and adaptation skills.
  • Uni Tree G1 Cold Weather Challenge: The Uni Tree G1 robot completed a 130,000-step walking challenge in -47°C temperatures, forming an Olympic logo with its path.
  • Inter Prior: An AI that helps robots learn to interact with objects in a natural and coordinated way through virtual training.

Speech Editing and Deepfake Concerns

A new AI technology allows for seamless editing of spoken content within videos, removing stutters, filler words, and even altering or inserting entire sentences without creating jarring “hard cuts.” The AI can condense longer passages into concise statements, demonstrating its ability to rephrase and restructure speech. However, this capability raises significant ethical concerns about the creation of deepfakes and the potential erosion of trust in video authenticity. As the presenter noted, “Pretty soon we're not going to be able to tell what's real and what's not real.” The tools are currently accessible via foul.ai and Pipio, with potential for future open-source release (though currently unconfirmed).

Conclusion

This week’s developments demonstrate a clear acceleration in AI capabilities, with both closed-source and open-source models pushing the boundaries of what’s possible. The increasing accessibility of powerful open-source models, coupled with advancements in video editing and speech manipulation, presents both exciting opportunities and significant ethical challenges. The rapid pace of innovation necessitates continued monitoring and discussion regarding the responsible development and deployment of these technologies.

Chat with this Video

AI-Powered

Load the transcript when you're ready to chat so the initial page stays lighter.

Ready to summarize another video?

Summarize YouTube Video