Key Concepts
- VAI's Open Source Model Series: A comprehensive family of models released by VAI, including language, multimodal understanding, image generation, and video generation models.
- GM 4.6 Series: VAI's latest flagship model series, demonstrating significant improvements in math and coding benchmarks.
- CC Bench: VAI's custom benchmark designed to test agent-style coding in real-world scenarios.
- Training Methodology: A multi-stage training process involving general pre-training, reasoning continuation training, mid-training, synthetic reasoning data, and long content/agent data.
- SLIDE Framework: VAI's reinforcement learning framework designed for efficient training of large language models, supporting both synchronous and decoupled modes.
- GM 4.5 V: VAI's latest visual understanding model supporting both image and video understanding.
- Agent Capability: The ability of models to control computers and interact with environments using tools like mouse and keyboard.
- Open Source Deployment: Methods for using VAI's models, including open-source weights with frameworks like Echelon and VLM, and cloud-based platforms like Z.AI.
VAI's Open Source Model Series and GM 4.6
VAI has a strong commitment to open-sourcing their work, starting with the G30B model in 2022. They have released a diverse family of models across various domains, visualized on a map with different colors representing model types: white for language (GM series), pink for multimodal understanding (GMV, formerly CodeVM), green for image generation, and yellow for video generation.
2025 is highlighted as VAI's "open source year," with the release of additional models like the GM4.0414 dense model (9B and 32B) and the GM4.5 MO series, their first MO (multimodal output) models. To date, VAI has released over 65 models, accumulating over 100 million downloads on platforms like Hugging Face and Monoscope. The community's engagement is evident with 1,500 community projects on GitHub related to GM or video models, fostering a vibrant, community-driven ecosystem.
GM 4.6: Performance and Benchmarks
GM 4.6 is presented as VAI's latest flagship model, showcasing significant advancements, particularly in math and coding benchmarks. It demonstrates a clear improvement over GM 4.5 and even surpasses some open-source models released concurrently, such as DeepSeek V3.2. Notably, GM 4.6 also outperforms commercial models like Claude 4 on several benchmarks, though a gap remains compared to Claude 4.5.
A key highlight is GM 4.6's performance on the Arena benchmark, which is considered closer to real user preferences. On Arena, GM 4.6 ranks as a top performer, sharing the number one spot with GPT-5 and Claude 4.5. It is the only open-source model achieving this ranking, a point of pride for VAI and a testament to the developers who tested and voted for it.
CC Bench: Agent-Style Coding Evaluation
VAI has developed its own dataset and benchmark called CC Bench to evaluate agent-style coding in real-world scenarios, moving beyond isolated problems. This platform is built on Cloud Code and features CC Bench version 1.1, which includes 22 new challenging coding tasks. The benchmark covers a wide range of applications, including front-end development, internal tool development, data analysis, and algorithm implementation.
In CC Bench, GM 4.6 shows a substantial leap over GM 4.5, achieving a 68.6% win rate against Claude 4, while significantly outperforming other open-source baselines. The benchmark records the full agent interaction, including query, planning, code generation, and execution, and is fully open-source for community access.
GM 4.6 Training and Data Strategy
The performance of GM 4.6 is attributed to its sophisticated training process, which involves several stages:
-
General Pre-training: The model is initially trained on approximately 15 trillion tokens of diverse data, including webpages, books, Wikipedia, and multilingual content. This stage aims to build a robust, all-around base model with a context window of 4,000 tokens.
-
Reasoning Continuation Training: This stage builds upon the base model by incorporating about 7 trillion tokens of additional code and reasoning data. This includes high-quality open-source reports and math/science problems requiring four-step reasoning.
-
Mid-training: The focus shifts to ripple-labeled code, encompassing multiple files, issues, and pull requests from projects. This data is packed into a single long context to teach the model to understand file dependencies, project structures, and changes within a project. The context window is expanded to 32,000 tokens, allowing the model to process key files of medium-sized repositories in one shot.
-
Synthetic Reasoning Data: Approximately 500 billion tokens of synthetic reasoning data are added, covering math, science, and algorithms with explicit thinking traces. This lays the groundwork for future agent behaviors like task decomposition, error reflection, and long-chain reasoning.
-
Long Content and Agent Data: Finally, about 100 billion tokens of long content and agent data are used. The context window is further pushed to 128,000 tokens for GM 4.6, with a specific mention of 200,000 tokens for the model. This enables the model to handle multiple documents, entire codebases, and very long conversations simultaneously. It also includes extensive agent queries, multi-step tool calls, search, and code execution, significantly improving the model's long-context and agent capabilities.
SLIDE: Reinforcement Learning Framework
VAI has developed SLIDE, an in-house reinforcement learning (RL) framework built on an Aston inference stack. SLIDE is designed to be efficient and adaptable to different task requirements.
-
Task-Specific Design: For short reasoning tasks like math or code completion, a synchronous training and inference setup is optimal, where training and inference occur on the same GPU. This allows for immediate sampling from the latest policy after each batch update, maximizing GPU memory and compute utilization.
-
Agent Task Optimization: For agent tasks, such as real software engineering, which involve many sequential steps (e.g., opening a browser, hitting an API), a decoupled and synchronized mode is employed. In this mode, workers interact directly with real environments, generate trajectories, and write them to a buffer. The training side then consumes this data at its own pace. This approach prevents slow tasks from bottlenecking the entire pipeline.
-
SLIDE Architecture: The framework features a meatron batch training engine (blue part) for data buffering and processing, and a high-throughput inference cluster (green part). A data buffer acts as a shared system, connecting training and agent environments.
-
Efficiency Optimizations: SLIDE incorporates efficiency optimizations, such as running the main chain in BF16 precision for stability. After each policy update, blockwise FP16 quantization is performed on the latest weights, and the FP16 version is sent through workers. This allows for higher throughput during training while maintaining BF16 precision, balancing accuracy and speed.
Reasoning and Data Quality in RL
VAI's RL training employs specific strategies to enhance performance:
-
Two-Stage Curriculum Learning: Instead of using a fixed dataset from start to finish, a two-stage curriculum is used. Stage one focuses on medium difficulty problems with varied rewards, allowing the model to strengthen. Stage two introduces extremely hard problems, where even occasional correct solutions are valuable. This approach, visualized by a blue curve, shows continuous improvement, outperforming a red curve that uses only median difficulty.
-
Single-Stage RL with Long Context: For models already trained with extensive Supervised Fine-Tuning (SFT) on 64,000 tokens, VAI found that multi-stage RL with shorter context windows can lead to forgetting long-context abilities. Therefore, they advocate for direct training with 64,000 tokens in a single stage, which clearly outperforms multi-stage approaches (red curve vs. blue curve).
-
Token-Weighted Loss for Code: For code generation, VAI uses a token-weighted loss (red curve) instead of a classic sequence loss (blue curve). The token-weighted loss averages loss over tokens, leading to faster and more stable convergence and reducing the generation of short, reward-gaming answers.
-
Data Quality over Quantity: In scientific reasoning, VAI found that a small, clean dataset of expert-verified, high-quality multiple-choice questions yields better performance than a mixed-quality dataset (red curve vs. blue curve on GPQA). This emphasizes that for scientific reasoning, data quality is paramount.
GM 4.5 V: Multimodal Understanding
GM 4.5 V is VAI's latest visual understanding model, excelling in grounding and image understanding benchmarks. It demonstrates clear advantages over other open-source models released around the same time.
-
Visual Input Processing: The model aims to preserve the original visual input as much as possible, accepting images at their native resolution and aspect ratio rather than forcing them into a fixed square. This is crucial for handling screenshots and long vertical images.
-
Video Understanding: For video, a time index token is inserted after each frame, indicating its temporal position. This helps the model understand temporal order and coherence, which is essential for action understanding and step-by-step processes.
-
GUI Agent Capability: GM 4.5 V also supports GUI agent capabilities, allowing it to control computers and interact with websites using mouse and keyboard inputs. This enables communication with browsers, computers, and mobile environments.
How to Use GM 4.6 and GM 4.5 V
VAI provides multiple avenues for users to leverage their models:
-
Open Source Weights: Users can download and use the open-source weights with frameworks like Echelon or VLM. Integration with third-party frameworks like Llama Factory and MS Swift is also supported, thanks to community contributions.
-
Cloud Deployment: For users without extensive GPU resources, VAI offers cloud-based solutions. The Z.AI website allows direct interaction with the models for tasks like writing code and generating presentations. A demo showcases using the model for Google searches with a single command.
-
GM Coding Playground: This platform connects GM models with tools and plugins (e.g., Cloud Code) to provide a powerful coding assistant experience. A demo video illustrates replacing a YOLO model in a Cocoa live app with GM 4.6.
Community Engagement and Resources
VAI actively engages with its community through various activities:
-
Regular Events: Both online and offline events are hosted, especially after new model releases, including AMAs on Reddit and on-site technology sharing sessions.
-
Important Links: Key resources are provided:
- Website: Z.AI for trying GM models.
- API: Available for programmatic access.
- Technical Reports: Detailed reports for GM 4.6 and GM 4.5 are available.
- Community: Discord link for joining discussions.
- GitHub: Link to open-source models with deployment instructions in the README.
Conclusion
VAI's GM 4.6 series represents a significant advancement in open-source AI, demonstrating state-of-the-art performance in key benchmarks, particularly in math and coding. The detailed training methodology, including extensive data curation and the innovative SLIDE RL framework, underpins this success. The GM 4.5 V model further expands VAI's multimodal capabilities. VAI's commitment to open-sourcing, coupled with robust community support and accessible deployment options, empowers developers and researchers to leverage these powerful models for a wide range of applications.
AI summaries can miss context or contain errors. Check important details against the original video.





