Key Concepts
- BU (Baidu) open-source LLMs: Text-to-text LLMs, VLMs (Vision Language Models), small dense models
- Model Variants: Base and Post-trained (PT) versions; PT is optimized for human preference
- Mixture of Experts (MoE): Architecture using multiple smaller models for efficiency
- Context Window: 128K tokens for all models
- Licensing: Apache 2.0 (allows commercial use)
- Integration: Hugging Face, OpenRouter (via Novita AI), Olama (unofficially), VLLM, BU's own platform
- KingBench: Test application used for demonstration
BU Open-Source LLMs Overview
Baidu (BU), a Chinese company, has released a suite of open-source language models, including text-to-text LLMs, vision language models (VLMs), and small dense models. All models feature a 128K context window and come in two variants: a base model and a post-trained (PT) model, with the PT version generally recommended for its optimization for human preference.
Text-to-Text LLMs
- 300B MoE Model: A mixture of experts model with 47 billion active parameters.
- 21B Model: A smaller, locally usable model with 21 billion total parameters and approximately 3 billion active parameters.
Vision Language Models (VLMs)
These models build upon the text-to-text LLMs, adding visual understanding capabilities.
- 424B Model: With 47B active parameters.
- 28B Model: With 3 billion active parameters.
Small Dense Model
- 0.3B (300 Million) Parameter Model: A simple, non-MoE dense model.
Availability and Licensing
All models are open-source and available on Hugging Face under the Apache 2.0 license, allowing for commercial use. They are also accessible via Baidu's chat platform for free. OpenRouter provides access to the models through Novita AI.
Benchmarks and Performance
- The 300B model outperforms Deepseek V3 and GPT-4.1 on benchmarks.
- The 21B model rivals the Quen 330B model, while being smaller and locally runnable.
- The VLMs add visual understanding with a relatively small parameter increase (120B for the larger model, 7B for the smaller).
Comparison with Other Models
| Model | Parameters | Active Parameters | Vision Capabilities | Local Runnability | | ------------------- | ----------------- | ----------------- | ------------------- | ----------------- | | BU 300B (Text) | 300B | 47B | No | No | | BU 21B (Text) | 21B | 3B | No | Yes | | BU 424B (VLM) | 424B | 47B | Yes | No | | BU 28B (VLM) | 28B | 3B | Yes | Yes | | Deepseek V3 | N/A | N/A | No | No | | Quen 330B | 330B | 3B | No | Difficult | | Metro Small | N/A | N/A | Yes | Yes (?) |
Using the Models
The models can be used via:
- Hugging Face integration: With VLLM or BU's deployment option.
- OpenRouter: Through Novita AI (bigger models, free variant expected).
- Kilo Code: Using the Ernie models (for free).
- Olama: Local model, better than Quen, matches or exceeds Devstrol
Integration with Development Tools (VS Code)
The video demonstrates integration with VS Code using Klein and R code.
- Klein: Configure via OpenRouter or the official BU API endpoint.
- R code: Similar configuration process to Klein.
Demonstration: Adding Dark/Light Theme Toggle to KingBench App
The video showcases the 400B model's ability to add a dark/light theme toggle to the KingBench app, a task complex even for Gemini. The smaller model couldn't achieve this, but is comparable to Devstrol in tool calling. The larger model sometimes outputs Chinese text, a potential inference issue.
Ninja Chat
An AI platform providing access to top AI models like GPT-4o, Claude 3.7 Sonnet, and Gemini 2.0 Flash for $11 per month. It features an AI playground for comparing model responses and a mind map generator. Use code king25 for 25% off any plan or king40yearly for 40% off annual subscriptions.
Conclusion
Baidu's open-source LLMs offer competitive performance, with the smaller models (21B text, 28B VLM) being particularly appealing for local use due to their performance and vision capabilities. The models are easy to integrate with development tools and offer a viable alternative to existing open and closed-source options. The 27B (likely referring to the 28B VLM) model is especially recommended for local development due to its balance of performance and vision capabilities.
AI summaries can miss context or contain errors. Check important details against the original video.