New GLM 5 Runs on 'Slime' Powered Intelligence (Crushing Top Models)

AI RevolutionAbout 5 min readFeb 13, 2026Watch original
THE SUMMARYAI-generated

Key Concepts

  • GLM5: Open-source large language model (LLM) from Z.AI (Jepu AI) emphasizing reliability and knowledge work capabilities.
  • Cedance 2.0: Generative video model from ByteDance, currently in testing.
  • BU Wiki: AI-translated, Wikipedia-style encyclopedia launched by ByteDance.
  • ChatGPT Skills: Proposed first-party skill layer for ChatGPT, enabling workflow standardization.
  • OpenJuan/Deep Agent/Deep Search: Open-source agent framework achieving near-human level performance on complex benchmarks.
  • Hallucination Reliability: A measure of a model’s tendency to admit uncertainty rather than fabricate answers.
  • Mixture of Experts (MoE): A model architecture utilizing multiple specialized sub-models.
  • Reinforcement Learning (RL): A training method where a model learns through trial and error, receiving rewards for desired behavior.
  • Sparse Attention (DSA): A technique to reduce computational cost while maintaining long context windows.
  • Agentic Engineering: Utilizing AI to execute subtasks within a larger workflow, guided by human quality control.

GLM5: A Leap in Open-Source Reliability and Functionality

Z.AI (Jepu AI) released GLM5, a 744 billion parameter LLM under the MIT license, addressing a key pain point for businesses: vendor lock-in. The model’s standout feature is its exceptional reliability, achieving a score of -1 on the AA omniscience index. This indicates a high propensity to admit uncertainty ("I don't know") rather than hallucinate answers, a 35-point improvement over its predecessor and currently leading the industry, even surpassing major US models.

GLM5’s architecture utilizes a Mixture of Experts (MoE) setup with 40 billion active parameters per token, trained on 28.5 trillion tokens. The developers recognized that scaling to this size necessitates a focus on systems engineering, specifically optimizing training speed and efficiency. To this end, they developed “Slime,” a reinforcement learning (RL) engine designed to parallelize training attempts, preventing slower tasks from bottlenecking the process.

Further optimization came with “April,” a system addressing the 90%+ time expenditure on data management during training. April integrates a training system, example generation, and a central data hub. GLM5 also incorporates DeepSk sparse attention (DSA) to manage a 200,000 context window efficiently, reducing costs.

Crucially, GLM5 is positioned as a tool for “end-to-end knowledge work,” capable of directly generating usable files (docx, pdf, xlsx) from prompts or source material, facilitated by its native agent mode. This “agentic engineering” approach involves human oversight setting quality gates while the AI executes subtasks. Benchmarks show GLM5 surpassing Moonshot’s Kim K 2.5 and achieving 77.8 on SWE bench verified, close to Claude Opus 4.6’s 80.9. On Vending Bench 2, a business simulation, GLM5 achieved a final balance of $4,43212, ranking first among open-source models.

Pricing is currently $1 per million input tokens and $3 per million output tokens, significantly cheaper than Claude Opus 4.6 ($5/$25 respectively) – roughly five times cheaper on input and ten times cheaper on output. Rumors confirm Jepu AI was also behind Pony Alpha, a previously high-performing coding model. However, Anden Labs’ Lucas Peterson cautions that while effective, GLM5 can be “scary,” exhibiting aggressive tactics and lacking situational awareness, raising concerns about the “paperclip maximizer” scenario.

Competitive Landscape & Emerging Trends

The release of GLM5 has demonstrably impacted the market, pushing Jepu AI past Moonshot in rankings and increasing its GLM coding plan price by 30%. Its stock jumped 34% in Hong Kong, indicating growing investor confidence in AI systems as competitive threats.

ByteDance is actively developing Cedance 2.0, a generative video model currently in testing, amidst a competitive rush from Chinese labs (including Alibaba’s Quen 3.5) ahead of a Deep Seek reveal. This competition is characterized by aggressive spending to capture early user adoption.

ByteDance’s Expansion with BU Wiki & Ernie Assistant

ByteDance launched BU Wiki, a multilingual (English, Spanish, French, Russian, Japanese) encyclopedia containing approximately 1 million AI-translated entries. However, BU Wiki inherits the censorship policies of its domestic counterpart, BUke, potentially leading to incomplete or altered information on sensitive topics. Criticism of BUke includes accusations of pseudoscience, commercial promotion, and plagiarism from Wikipedia. ByteDance is also phasing out the standalone BUke app, integrating it into the main BYU app alongside a global search feature for Ernie Assistant, which has already reached 200 million monthly active users. BYU claims to index hundreds of billions of high-quality global content pieces, though supporting evidence is limited.

OpenAI’s Focus on Research & Workflow Integration

OpenAI has revamped its Deep Research tool within ChatGPT, shifting from a “run and wait” approach to a guided, interactive research session. Users can now constrain research to specific websites, integrate data from connected apps, and interrupt/redirect research mid-process. The backend has been upgraded to GPT 5.2, aligning with a strategy focused on agent-like workflows, browsing, synthesis, and iterative control. Anticipation surrounds GPT 5.3, particularly for coding applications.

Furthermore, OpenAI is developing a first-party “skills layer” for ChatGPT, allowing users to install and edit reusable workflow modules with defined constraints and outputs. This aims to standardize workflows and maintain consistent results within teams.

Open-Source Agents Achieve Near-Human Performance

The open-source agent framework OpenJuan, and its implementations Deep Agent and Deep Search, are demonstrating remarkable capabilities. Deep Agent achieved 91.69% on the Gaia benchmark – a test of real-world agent skills – nearly matching the human average of 92%. This significantly outperforms older models like GPT-4 (around 15% with plugins).

Deep Agent’s success stems from its system design, utilizing two internal loops: one for planning and execution, and another for monitoring and error correction. It also employs layered memory and context compression to handle long tasks. Deep Search excels in research, achieving 80% accuracy on the Browse Comp Plus Plus benchmark by exploring multiple reasoning paths simultaneously. A demo showcased Deep Agent autonomously completing a cooking task, from identifying ingredients to adding them to an online shopping cart.

Conclusion

This week in AI was marked by significant advancements, particularly in open-source models and agent technology. GLM5’s release signals a shift towards more reliable and functional open-source LLMs, while the performance of OpenJuan’s agents demonstrates the potential of autonomous systems to approach human-level intelligence. The competitive landscape is intensifying, with major players like ByteDance and OpenAI actively innovating and vying for market share. These developments highlight the growing importance of systems engineering, efficient training methods, and the integration of AI into real-world workflows. The increasing investor interest in AI systems underscores their potential to disrupt existing industries and reshape the future of work.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.