UX Design Principles for Semi Autonomous Multi Agent Systems — Victor Dibia, Microsoft

AI EngineerAbout 7 min readJul 22, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Multi-Agent Systems: Systems where multiple AI agents collaborate to solve a problem.
  • Autogen: An open-source framework for building multi-agent applications.
  • BlenderLM: A multi-agent system built from scratch for enabling 3D tasks in Blender.
  • Fixed Deterministic Workflow: A workflow with pre-defined steps and solutions.
  • Autonomous Exploratory Systems: Systems where the LLM drives the flow of control, takes actions, observes results, and makes progress.
  • Capability Discovery: Identifying and showcasing the tasks that an agent can perform with high reliability.
  • Observability and Provenance: Providing detailed activity logs and debugging tools to help users understand the agent's actions.
  • Interruptibility: Designing the system to allow users to pause, checkpoint, rollback, and resume the agent's processes.
  • Cost-Aware Delegation: Ensuring agents can quantify the risk or cost of their actions and delegate to users when necessary.
  • Eval Driven Design: Defining evaluation metrics and building baselines to iteratively improve agents.

I. Introduction

  • Victor Dia, a Principal Research Software Engineer at Microsoft Research, discusses his work on human-AI experiences, particularly focusing on scenarios where humans and AI agents collaborate to solve problems.
  • He highlights GitHub Copilot as a successful example of an AI model assisting developers at scale and introduces Autogen, an open-source multi-agent framework, and Autogen Studio, a local developer tool for building multi-agent workflows.
  • He shares a brief history of his journey into agents, starting with Project LA.

II. Project LA: An Early Agentic Workflow

  • In August 2022, before the widespread adoption of Chat GPT, Victor worked on Project LA.
  • Functionality: It allowed users to drag data (CSV or JSON files) into a web interface and automatically performed data summarization, question answering, code generation, execution, post-processing, error recovery, and visualization generation.
  • Components: It comprised four main categories: summarization, goal exploration, visualization generation, and a built-in code interpreter.
  • Technology: The first version utilized OpenAI's Da Vinci Codeex models, initially resulting in a 20% error rate, which later decreased to 1.5% with the introduction of the GPT-3.5 Turbo models.
  • Impact: Victor notes that Project LA demonstrated the potential of these applications and similar capabilities can now be seen in various Microsoft products and beyond.

III. Autogen and Autogen Studio

  • Inspired by Project LA, Victor and his colleagues began exploring the development of multi-agent applications where agents exchange messages and self-organize to solve problems.
  • Autogen: A framework for building multi-agent applications with significant usage (5k stars on GitHub).
  • Autogen Studio: A low-code tool that allows users to compose multiple agents into teams, defining their models and tools to build multi-agent workflows through a web interface.

IV. BlenderLM: A Multi-Agent System for 3D Tasks

  • Instead of presenting existing Autogen Studio capabilities, Victor chooses to demonstrate a live build of BlenderLM, a multi-agent system designed to facilitate 3D tasks in Blender.
  • Purpose: BlenderLM aims to translate natural language instructions into 3D artifacts within Blender.
  • Motivation: Inspired by Victor's personal experience learning Blender, where completing the "donut tutorial" required extensive time (40-50 hours).

V. Multi-Agent System Design Considerations

  • Fixed Deterministic vs. Autonomous Exploratory Systems: The presentation contrasts two approaches to multi-agent systems.
    • Fixed Deterministic Workflows: Suitable for problems with known solutions, allowing for reliable systems using function calling and structured output.
    • Autonomous Exploratory Systems: Necessary when the solution is not known in advance, requiring the LLM to drive the flow of control through tool use, action taking, result inspection, and observation.
  • Characteristics of Autonomous Exploratory Systems:
    1. Autonomy: Ability to address multiple tasks.
    2. Action-Taking with Side Effects: Actions can have unexpected results that the system must handle.
    3. Exploration of Complex Tasks: Ability to break down tasks into steps and execute them over extended periods.

VI. BlenderLM Demo

  • A live demo of BlenderLM is presented.
  • Interface: The web application connects to a Blender instance via a WebSocket.
  • Tools: The interface provides a set of fixed tools (buttons) that the developer can use directly, such as clearing the scene.
  • Functionality: Users can select pre-determined examples or input natural language instructions (e.g., "Create two balls with a shiny glossy silver finish").
  • Real-time Updates: The system streams activities in real-time, showing the analysis, plan, and execution steps on the UI and in the Blender interface.
  • Agent Interaction: The demo showcases the planning agent breaking down the task into steps and the system using a verification loop to assess progress using an LLM.
  • Result: The system successfully creates two shiny, silver-colored balls in the Blender scene.

VII. Building BlenderLM: A Step-by-Step Process

  • The development process of BlenderLM is outlined in a series of steps:
    1. Define the Goal: Translate natural language tasks into 3D artifacts.
    2. Create a Baseline: Develop a simple script that adds a single cube to the scene (the "hello world" of Blender). This baseline involves building a Blender add-on and a client library for socket connections, which are essential for rapid prototyping and testing.
    3. Build Out Tools: Define the tools needed by the agent, divided into:
      • Task-Specific Tools: Directly create Blender objects.
      • General-Purpose Tools: Execute arbitrary code generated by the LLM to drive Blender's capabilities.
      • Note: Agent's quality depends on the tools it has.
    4. Define a Test Bed (Eval):
      • V1: Jupyter Notebook: Initially, code is tested within a Jupyter notebook.
      • V2: Interactive Web UI: Then, a fully interactive web UI (as demoed) is created.
      • V3: Automated Test Suite: Ultimately, an automated test suite with metrics and a full evaluation harness is required.
    5. Build the Agent:
      • Base Agent Loop: Create a basic LLM loop with function calls.
      • Verifier Agent: Add a verifier agent that takes snapshots of the scene content, using an LLM to predict progress and task completion to guide the process.
      • Planner Agent: Implement a planner agent to break down incoming tasks into atomic steps.

VIII. Design Principles for Multi-Agent Assistant User Experience

  • Victor emphasizes that the following design principles are not exhaustive but offer a starting point for improving multi-agent systems:
    1. Capability Discovery: Itemize and showcase the tasks the agent can reliably perform.
    2. Observability and Provenance: Stream activity logs and provide debugging tools (e.g., token usage, time taken) to help users understand the agent's actions.
    3. Interruptibility: Enable users to pause, checkpoint, rollback, and resume agent processes.
    4. Cost-Aware Delegation: Quantify the risk or cost of agent actions, delegating potentially harmful or expensive actions to users for approval.

IX. Key Takeaways

  1. Know When to Use a Multi-Agent Approach:
    • Multi-agent systems are not always the right solution.
    • Increased autonomy also increases the surface for error.
    • Carefully assess the problem space to ensure a multi-agent system is the appropriate tool.
    • A small percentage of tasks truly benefit from a multi-agent approach.
    • Five-Step Framework: to determine if a task might benefit from a multi-agent approach:
      • Does the task benefit from planning?
      • Can the task be broken into multiple perspectives or personas?
      • Does the task require consuming or processing extensive context?
      • Does the task require consuming or processing extensive context?
      • Adaptive solutions.
  2. Eval-Driven Design:
    • Define the task and evaluation metrics before building the agent.
    • Build a baseline without agents.
    • Iteratively improve the agents.
    • Academic benchmarks are great, but you should build evals that are tuned to your specific task.
  3. Design Principles:
    • Ensure users can discover the ideal tasks for the system.
    • Provide user-facing observability traces.
    • Ensure agents are interruptible.
    • Ensure agents can quantify action costs and delegate to users as needed.
  4. Avoid Building Everything from Scratch:
    • Leverage frameworks to streamline development.

X. Conclusion

  • Victor concludes by providing further reading references, including papers on Autogen Studio, magentic UI, and challenges in human-AI communication. He also mentions his upcoming book, which dedicates a chapter to design principles. The code for BlenderLM is also made available. The presentation emphasizes the importance of careful planning, iterative development, and user-centric design when building multi-agent systems.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.