THE SUMMARYAI-generated
Key Concepts
- Multi-Agent Systems: Systems where multiple AI agents collaborate to solve a problem.
- Autogen: An open-source framework for building multi-agent applications.
- BlenderLM: A multi-agent system built from scratch for enabling 3D tasks in Blender.
- Fixed Deterministic Workflow: A workflow with pre-defined steps and solutions.
- Autonomous Exploratory Systems: Systems where the LLM drives the flow of control, takes actions, observes results, and makes progress.
- Capability Discovery: Identifying and showcasing the tasks that an agent can perform with high reliability.
- Observability and Provenance: Providing detailed activity logs and debugging tools to help users understand the agent's actions.
- Interruptibility: Designing the system to allow users to pause, checkpoint, rollback, and resume the agent's processes.
- Cost-Aware Delegation: Ensuring agents can quantify the risk or cost of their actions and delegate to users when necessary.
- Eval Driven Design: Defining evaluation metrics and building baselines to iteratively improve agents.
I. Introduction
- Victor Dia, a Principal Research Software Engineer at Microsoft Research, discusses his work on human-AI experiences, particularly focusing on scenarios where humans and AI agents collaborate to solve problems.
- He highlights GitHub Copilot as a successful example of an AI model assisting developers at scale and introduces Autogen, an open-source multi-agent framework, and Autogen Studio, a local developer tool for building multi-agent workflows.
- He shares a brief history of his journey into agents, starting with Project LA.
II. Project LA: An Early Agentic Workflow
- In August 2022, before the widespread adoption of Chat GPT, Victor worked on Project LA.
- Functionality: It allowed users to drag data (CSV or JSON files) into a web interface and automatically performed data summarization, question answering, code generation, execution, post-processing, error recovery, and visualization generation.
- Components: It comprised four main categories: summarization, goal exploration, visualization generation, and a built-in code interpreter.
- Technology: The first version utilized OpenAI's Da Vinci Codeex models, initially resulting in a 20% error rate, which later decreased to 1.5% with the introduction of the GPT-3.5 Turbo models.
- Impact: Victor notes that Project LA demonstrated the potential of these applications and similar capabilities can now be seen in various Microsoft products and beyond.
III. Autogen and Autogen Studio
- Inspired by Project LA, Victor and his colleagues began exploring the development of multi-agent applications where agents exchange messages and self-organize to solve problems.
- Autogen: A framework for building multi-agent applications with significant usage (5k stars on GitHub).
- Autogen Studio: A low-code tool that allows users to compose multiple agents into teams, defining their models and tools to build multi-agent workflows through a web interface.
IV. BlenderLM: A Multi-Agent System for 3D Tasks
- Instead of presenting existing Autogen Studio capabilities, Victor chooses to demonstrate a live build of BlenderLM, a multi-agent system designed to facilitate 3D tasks in Blender.
- Purpose: BlenderLM aims to translate natural language instructions into 3D artifacts within Blender.
- Motivation: Inspired by Victor's personal experience learning Blender, where completing the "donut tutorial" required extensive time (40-50 hours).
V. Multi-Agent System Design Considerations
- Fixed Deterministic vs. Autonomous Exploratory Systems: The presentation contrasts two approaches to multi-agent systems.
- Fixed Deterministic Workflows: Suitable for problems with known solutions, allowing for reliable systems using function calling and structured output.
- Autonomous Exploratory Systems: Necessary when the solution is not known in advance, requiring the LLM to drive the flow of control through tool use, action taking, result inspection, and observation.
- Characteristics of Autonomous Exploratory Systems:
- Autonomy: Ability to address multiple tasks.
- Action-Taking with Side Effects: Actions can have unexpected results that the system must handle.
- Exploration of Complex Tasks: Ability to break down tasks into steps and execute them over extended periods.
VI. BlenderLM Demo
- A live demo of BlenderLM is presented.
- Interface: The web application connects to a Blender instance via a WebSocket.
- Tools: The interface provides a set of fixed tools (buttons) that the developer can use directly, such as clearing the scene.
- Functionality: Users can select pre-determined examples or input natural language instructions (e.g., "Create two balls with a shiny glossy silver finish").
- Real-time Updates: The system streams activities in real-time, showing the analysis, plan, and execution steps on the UI and in the Blender interface.
- Agent Interaction: The demo showcases the planning agent breaking down the task into steps and the system using a verification loop to assess progress using an LLM.
- Result: The system successfully creates two shiny, silver-colored balls in the Blender scene.
VII. Building BlenderLM: A Step-by-Step Process
- The development process of BlenderLM is outlined in a series of steps:
- Define the Goal: Translate natural language tasks into 3D artifacts.
- Create a Baseline: Develop a simple script that adds a single cube to the scene (the "hello world" of Blender). This baseline involves building a Blender add-on and a client library for socket connections, which are essential for rapid prototyping and testing.
- Build Out Tools: Define the tools needed by the agent, divided into:
- Task-Specific Tools: Directly create Blender objects.
- General-Purpose Tools: Execute arbitrary code generated by the LLM to drive Blender's capabilities.
- Note: Agent's quality depends on the tools it has.
- Define a Test Bed (Eval):
- V1: Jupyter Notebook: Initially, code is tested within a Jupyter notebook.
- V2: Interactive Web UI: Then, a fully interactive web UI (as demoed) is created.
- V3: Automated Test Suite: Ultimately, an automated test suite with metrics and a full evaluation harness is required.
- Build the Agent:
- Base Agent Loop: Create a basic LLM loop with function calls.
- Verifier Agent: Add a verifier agent that takes snapshots of the scene content, using an LLM to predict progress and task completion to guide the process.
- Planner Agent: Implement a planner agent to break down incoming tasks into atomic steps.
VIII. Design Principles for Multi-Agent Assistant User Experience
- Victor emphasizes that the following design principles are not exhaustive but offer a starting point for improving multi-agent systems:
- Capability Discovery: Itemize and showcase the tasks the agent can reliably perform.
- Observability and Provenance: Stream activity logs and provide debugging tools (e.g., token usage, time taken) to help users understand the agent's actions.
- Interruptibility: Enable users to pause, checkpoint, rollback, and resume agent processes.
- Cost-Aware Delegation: Quantify the risk or cost of agent actions, delegating potentially harmful or expensive actions to users for approval.
IX. Key Takeaways
- Know When to Use a Multi-Agent Approach:
- Multi-agent systems are not always the right solution.
- Increased autonomy also increases the surface for error.
- Carefully assess the problem space to ensure a multi-agent system is the appropriate tool.
- A small percentage of tasks truly benefit from a multi-agent approach.
- Five-Step Framework: to determine if a task might benefit from a multi-agent approach:
- Does the task benefit from planning?
- Can the task be broken into multiple perspectives or personas?
- Does the task require consuming or processing extensive context?
- Does the task require consuming or processing extensive context?
- Adaptive solutions.
- Eval-Driven Design:
- Define the task and evaluation metrics before building the agent.
- Build a baseline without agents.
- Iteratively improve the agents.
- Academic benchmarks are great, but you should build evals that are tuned to your specific task.
- Design Principles:
- Ensure users can discover the ideal tasks for the system.
- Provide user-facing observability traces.
- Ensure agents are interruptible.
- Ensure agents can quantify action costs and delegate to users as needed.
- Avoid Building Everything from Scratch:
- Leverage frameworks to streamline development.
X. Conclusion
- Victor concludes by providing further reading references, including papers on Autogen Studio, magentic UI, and challenges in human-AI communication. He also mentions his upcoming book, which dedicates a chapter to design principles. The code for BlenderLM is also made available. The presentation emphasizes the importance of careful planning, iterative development, and user-centric design when building multi-agent systems.
AI summaries can miss context or contain errors. Check important details against the original video.