From Mixture of Experts to Mixture of Agents with Super Fast Inference - Daniel Kim & Daria Soboleva

AI EngineerAbout 5 min readJun 27, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Mixture of Experts (MoE)
  • Mixture of Agents (MoA)
  • Large Language Models (LLMs)
  • Inference Time Compute
  • Hardware Acceleration
  • Prompt Engineering
  • Cerebras Architecture
  • Model Scaling

API Key and Workshop Overview

The workshop focuses on building a "Mixture of Agents" (MoA) application. Participants need to obtain a free API key from the provided QR code to access Cerebras' cloud services. The workshop agenda includes:

  1. Introduction to Mixture of Experts (MoE) and its significance.
  2. Hands-on session to build a MoA application.
  3. Q&A session about Cerebras and its models.

Mixture of Experts (MoE) Explained

  • Problem: Scaling LLMs requires architectural improvements in addition to data improvements.
  • Solution: MoE replaces the monolithic feed-forward network in a transformer architecture with multiple specialized feed-forward networks called "experts."
  • Mechanism: A "router" network decides which expert to select for a particular token.
  • Benefits:
    • Increases model parameters without proportionally increasing inference time.
    • Allows for specialization of experts (e.g., math, biology).
    • Industry standard for scaling models (used by OpenAI, Anthropic).

Inference Time Compute and Mixture of Agents (MoA)

  • Inference Time Compute: Focuses on improving model performance through more compute after pre-training, especially as data availability plateaus.
  • Example: Solving complex math problems requires reasoning, which is computationally expensive for monolithic models.
    • GPT-4 (non-reasoning): 45 seconds, wrong answer.
    • GPT-3 (reasoning): 293 seconds, correct answer.
  • MoA Approach: Leverages multiple LLMs (agents) with custom system prompts to solve problems collectively.
    • Inspired by MoE architecture.
    • A final model combines the answers from individual agents.
    • Can outperform frontier models on certain benchmarks.
  • Example Startup: ninjate.ai: Solves the same math problem in 7.4 seconds using MoA.
    • Planning agent generates proposals.
    • Critique agent evaluates proposals.
    • Summarization agent combines top answers into a final answer.
    • Process involves multiple LLM calls (e.g., 32 calls, 500,000 tokens).

Cerebras Hardware Advantage

  • GPU Bottleneck: GPUs have limited cores (e.g., 17,000 in H100) and rely on off-chip memory, creating bottlenecks due to memory transfer.
  • Cerebras Solution:
    • 900,000 cores on a single chip.
    • Each core has direct access to its own memory.
    • Weights are stored locally, eliminating the need to load them from external memory.
    • Scales linearly with larger models.
    • Only activations are transferred between chips.
  • Performance: Cerebras is significantly faster than GPUs for large models (e.g., 15.5x faster for Llama 3 70B inference).

Hands-on Workshop: Building a MoA Application

  • Goal: Specialize agents to solve specific parts of a complex problem, leading to better results with less prompting.
  • Process:
    1. Obtain API key.
    2. Clone the GitHub repository.
    3. Deploy the app via Streamlit or locally.
    4. Configure agents with custom prompts.
  • UI Components:
    • Summarization agent (aggregates results).
    • Agent management (create, modify, delete agents).
    • Layers (sequential processing of agents).

AI Configuration Challenge

  • Task: Configure a MoA system to generate Python code that scores the maximum 120 points in an automated grader.
  • Role: Become an AI prompt engineer and system architect.
  • Parameters to Adjust:
    • Main model.
    • Number of cycles (layers).
    • Temperature.
    • System prompt.
    • Individual layer prompts.
  • Challenge Details:
    • Create a function called calculate_user_matrix.
    • Fix bugs and optimize the function.
    • Use agents specialized in bug detection, performance optimization, and overall code quality.
  • Baseline Function: A buggy and unoptimized Python function is provided as a starting point.

Q&A Highlights

  • AutoML for Prompts: While manual prompt engineering is currently emphasized, the future involves automating this process, similar to codegen startups like Devon.
  • Hardware Availability: Cerebras has multiple data centers in the US and plans to expand globally, including France and Canada.
  • Model Onboarding: The time to onboard a new model depends on the similarity of its architecture to existing supported models and the availability of necessary kernels.
  • Power Consumption: Cerebras claims to have lower power consumption compared to Nvidia GPUs for equivalent workloads.
  • MoA Tradeoffs: Creating too many agents can lead to redundancy and increased computation time without improving the final solution.
  • Fine-tuned Models: Cerebras supports bringing your own fine-tuned models for enterprise clients and is working on supporting LoRA fine-tuned models.
  • Diffusion Models: While not currently a focus, Cerebras is exploring diffusion models and has seen promising internal demos.
  • Custom Architectures: Cerebras collaborates with customers to support custom architectures by developing the necessary kernels.
  • Real-time APIs: Cerebras is considering real-time APIs and multimodal models, with some image-based queries already running on Cerebras hardware through Mistral.
  • Models Engineered for Cerebras: While most models are ported, Cerebras offers advantages for unstructured sparsity algorithms and is working with companies like Mistral to design models specifically for Cerebras hardware.
  • Model Sizes: Cerebras can support a wide range of model sizes, scaling linearly by adding more chips.

Synthesis/Conclusion

The workshop introduces the concept of Mixture of Agents (MoA) as a method to improve LLM performance by leveraging multiple specialized models. Cerebras' hardware architecture, with its massive core count and distributed memory, offers significant advantages for running MoA applications and achieving faster inference times. The hands-on session allows participants to experiment with prompt engineering and system architecture to optimize code generation, highlighting the potential of MoA for solving complex problems. The Q&A session addresses various aspects of Cerebras' technology, including hardware availability, model onboarding, power consumption, and future directions.

AI summaries can miss context or contain errors. Check important details against the original video.

MAKE IT YOURS

Read. Remember. Reuse.

Free tools

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.