Why they feel dumber? ... its not the model
By Prompt Engineering
Key Concepts
- Open-weight models: Large language models whose weights are publicly available.
- Inference provider: A service or platform that hosts and runs AI models, making them accessible via APIs.
- State-of-the-art (SOTA) results: The best performance achieved on a particular task or benchmark.
- Context length: The maximum amount of text (tokens) a model can process at once.
- Floating-point precision: The level of accuracy used to represent numbers in a computer, affecting model size and performance (e.g., 16-bit, 8-bit, 4-bit).
- Quantization: The process of reducing the precision of model weights, typically to decrease model size and inference cost, but potentially impacting accuracy.
- Academic benchmarks: Standardized tests used to evaluate model performance on specific tasks (e.g., MMLU).
- Real-world use cases: Practical applications of AI models.
- Agentic systems: AI systems designed to perform tasks autonomously, often involving decision-making and tool usage.
- Tool calling/Function calling: The ability of an LLM to identify when a specific external tool or function is needed to fulfill a user's request and to generate the correct parameters for that tool.
- Prompt template: A predefined structure for inputting prompts to an LLM, which can affect its output.
- Sampling configurations: Parameters that control how an LLM generates text, influencing its creativity and coherence.
- Backend as a Service (BaaS): A cloud service model that provides pre-built backend functionalities (like authentication, databases, storage) accessible via APIs, simplifying development.
- Schema validation errors: Errors that occur when the data provided to a tool or function does not match the expected structure or format.
- Euclidean distance: A mathematical measure of the straight-line distance between two points in a multi-dimensional space, used here to compare metric values.
Performance Discrepancies in Open-Weight Models and API Providers
The video highlights a critical issue in the adoption of large open-weight AI models: the significant performance variations observed when using different third-party API providers to host these models, even when the model weights themselves are identical. This discrepancy can lead to users experiencing performance far below advertised state-of-the-art (SOTA) results.
Key Points:
- The Problem: Users often cannot run massive models (e.g., 500 billion or 1 trillion parameters) locally and rely on external API providers. However, the performance achieved through these providers can be substantially different from what is expected or benchmarked by the model creators.
- Root Cause: The variation is not due to the model weights themselves but rather how the API providers implement and host the models.
Evidence and Examples
-
OpenRouter and GLM 4.6:
- OpenRouter lists various models, including GLM 4.6, with different providers.
- These providers offer varying context lengths, price points, and sometimes different floating-point precisions.
- Crucially, the video points out that the implementation by these providers leads to performance differences not explicitly stated on the platform.
-
GPT-Neo OSS and SemiAnalysis Benchmarks:
- When GPT-Neo OSS was released, SemiAnalysis conducted independent benchmarks on various third-party providers.
- Data: For the MMLU benchmark, performance varied significantly: the best provider achieved 93%, while the worst scored 80%, a 13-point difference. This demonstrates that the choice of API provider can critically impact performance on specific tasks.
-
Kimi.ai and Agentic Coding:
- Kimi.ai is designed for agentic coding and competitive coding.
- They observed significant differences in tool call performance across various open-source solutions and vendors.
- Quote: Kimi.ai states, "When selecting a provider, user often prioritize lower latency and cost but may overlook more subtle yet critical differences in model accuracy." They also note that this leads to inconsistent performance, which users wrongly attribute to the model creator.
-
Kimi.ai Vendor Verifier Benchmark:
- To address this, Kimi.ai developed the "Kimi.ai vendor verifier," a benchmark run regularly to assess API providers' performance on specific tasks.
- Focus: This benchmark is specifically designed for tool call or function calling capabilities, which are crucial for agentic systems.
- Shocking Results: The benchmark revealed drastic performance differences. The official first-party API (Moonshot) performed well, but some other providers showed performance as low as 50% compared to the official implementation.
- Implication: This underscores that cost and speed are not the only factors; performance accuracy is equally vital.
-
SemiAnalysis Inference Max:
- SemiAnalysis also released "Inference Max," a benchmark focusing on GPU performance variations.
Factors Contributing to Performance Variations
The video details several technical reasons why API providers can achieve different performance levels with the same model weights:
-
Prompt Template:
- Early Issue: In the past, different prompt templates for open-weight models led to significant performance degradation if the wrong template was used.
- Current Relevance: While more standardized now, OpenAI's recent release of its "harmony prompt template" with OSS models caused issues for some providers in implementing it correctly, leading to performance differences.
-
Quantization:
- Definition: Quantization involves reducing the floating-point precision of model weights (e.g., from 16-bit to 8-bit or 4-bit).
- Impact: Using lower precision (e.g., 8-bit or 4-bit) by some vendors compared to the first-party provider's 16-bit can lead to substantially different performance. 4-bit quantization can be particularly detrimental, especially for smaller models.
-
Configs and Sampling:
- Model Creator Guidance: Open-weight model creators often provide optimal configurations and sampling parameters.
- Provider Implementation: If API providers deviate from these recommendations or misconfigure settings, performance will suffer.
- Hosting Frameworks: Different hosting frameworks like vLLM or Transformers, or even specific implementations like llama.cpp, can introduce variations. Some might be optimized for throughput, while others might have misconfigurations.
Understanding Tool Calls and Agentic Systems
The Kimi.ai benchmark's focus on tool calls is explained in detail:
-
Anatomy of a Tool Call:
- Decision: The LLM first determines if a tool is needed for a query or if it can answer directly.
- Tool Selection: If a tool is required, the LLM selects the appropriate tool based on available descriptions.
- Schema Input: The LLM must then pass the correct schema (data structure) as input to the selected tool.
- Tool Execution: The system calls the tool with the provided input.
- Response Processing: The tool generates a response, which is then processed by the system.
- LLM Output: The LLM receives the tool's output and uses it to formulate its final response to the user.
-
Agentic Systems Architecture:
- Agentic systems typically involve a backend and a frontend.
- The backend often requires various services: authentication, database connectors (e.g., for storing information), storage services, and real-time analytics.
- Developers traditionally had to manage connectors for each service.
Backend as a Service (BaaS) and Superbase
- Emerging Trend: A new class of tools, Backend as a Service (BaaS), simplifies backend development.
- Example: Superbase (a sponsor of the video) is a prime example.
- Benefits: BaaS platforms provide a single API endpoint to interact with multiple backend functionalities, eliminating the need for individual connectors.
- Superbase Features:
- Offers authentication, database (built on PostgreSQL, a familiar technology), storage, and real-time analytics via a single API.
- Includes its own MCP server, which is particularly useful for agents interacting with Superbase functionalities.
- Relevance to Benchmarking: The Kimi.ai vendor verifier benchmark utilizes tools that interact with MCP servers, assessing agents' ability to discover and call these tools.
Metrics Used in the Kimi.ai Vendor Verifier Benchmark
The video outlines key metrics used to evaluate tool call performance, offering insights for building custom benchmarks:
finish_reasonisstopped: The number of responses where the model decided to generate a final response.- Actual Tool Calls Made: The total number of times the model attempted to use a tool.
- Schema Creation Errors: The number of instances where errors occurred in generating the schema for a tool call. This is a critical indicator of the model's ability to correctly format inputs for tools.
- Successful Tool Calls: The number of times the model successfully selected the correct tool, passed a valid schema, and received a response.
- Euclidean Distance: The Euclidean distance between the metric values of a provider and those of the official Moonshot AI. This provides a quantitative measure of deviation from the baseline.
Important Note: These metrics focus on the validity of tool calls, not the output of the tools themselves, as the latter is handled by the external system.
Benchmark Findings and Implications
- Failure Patterns: Models that perform poorly often make more tool calls or require more conversational turns. They also exhibit a higher rate of schema generation errors.
- Specific Observations:
- Grok showed zero schema validation errors, which is noted as "pretty amazing."
- Models hosted on Moonshot also had no schema validation errors.
- Other providers showed validation errors, suggesting potential issues like quantization or improper configuration.
- Call for Regular Benchmarking: Kimi.ai plans to run its benchmark regularly. The video advocates for other model creators to do the same.
- Benefits of Benchmarking:
- For Model Creators: Demonstrates that model weights are sound and highlights issues with hosting services.
- For API Providers: Creates pressure to ensure proper hosting and optimization.
- Business Opportunity: There's a potential business opportunity in creating such benchmarks, even for proprietary models, to track performance changes and identify issues (e.g., with cloud APIs or specific model versions).
Conclusion
The video strongly emphasizes that when evaluating and deploying open-weight LLMs, the choice of inference provider is as crucial as the model itself. Performance variations are significant and stem from technical implementation details like prompt templating, quantization, and sampling configurations. Benchmarks like the Kimi.ai vendor verifier are essential for transparency and accountability, helping users make informed decisions and pushing providers to deliver consistent, high-quality performance. The rise of BaaS solutions like Superbase further simplifies the integration of AI agents with backend services, making robust tool calling capabilities even more critical.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

Why Does This Guy Appear In Kids Videos?
sphynx

TIC en las Organizaciones - Electiva Complementaria II Unisimon
Julieth Güell S

How to Tame Your Advice Monster | Michael Bungay Stanier | TED
TED

Margaret Heffernan: Why it's time to forget the pecking order at work
TED

The importance of psychological safety: Amy Edmondson
The King's Fund

What Is Psychological Safety?
Harvard Business Review

13-Conflict Management: Listening in Conflict
Deliberate Development