THE SUMMARYAI-generated
Key Concepts
- Agentic Systems: Systems that automate tasks traditionally performed by humans, particularly in software engineering.
- Multi-Agent Systems: Systems composed of multiple interacting agents, each specializing in a specific task or area.
- LLMs (Large Language Models): The foundational technology used to build agents, providing natural language understanding and generation capabilities.
- Tools: External functions or APIs that LLMs can use to interact with the environment (e.g., querying databases, accessing logs, running commands).
- Feedback Loops: Mechanisms that allow agents to evaluate their own performance and refine their actions iteratively.
- Planning: The ability of an agent to reason about a task, break it down into steps, and execute those steps strategically.
- Memory: The ability of an agent to store and retrieve information about past experiences, allowing it to learn and improve over time.
- Guardrails: Constraints or limitations imposed on agents to prevent them from taking unintended or harmful actions.
- Context Window: The amount of text that an LLM can process at one time, which limits the complexity of tasks it can perform.
- Token Efficiency: Strategies for minimizing the number of tokens used by an LLM, reducing cost and improving performance.
How Agentic Systems for Software Engineering Look
- Goal: To build systems that automate the work of human software engineers.
- Complexity: Software engineering involves more than just coding; it includes maintaining production systems, ensuring reliability, and managing infrastructure.
- Current State: Software engineers use a variety of tools for coding, infrastructure, observability, and telemetry.
- Human as Glue: Humans currently act as the glue between these tools, combining inputs and outputs to perform tasks.
- Agent-Assisted: Agents can perform tasks within specific categories, with humans acting as the glue between agents.
- Multi-Agent Systems: The future involves multi-agent systems that take input from humans, interact with other agents, iterate on plans, and provide comprehensive solutions.
- Challenges:
- Infinite possibilities in data.
- Dynamic environments and dependencies.
- Tribal or institutional knowledge.
- Constantly changing environments.
Demo of an Agentic System (Resolve)
- Multi-Agent System: Agents understand different types of tools in a software engineering environment (infrastructure, metrics, events, code).
- Scenario 1: UI Slowdown:
- A user reports that the UI is slow.
- The agent takes this high-level directive and maps it to a set of tools.
- It plans how to solve the problem and maps it to the front-end service.
- The agent looks across all available tools to find the cause.
- The agent identifies spikes in UI front-end latency and a temporary flood of product lookup order processing errors.
- The agent is asked to find the root cause.
- The agent identifies that a feature flag was intentionally enabled at 8 a.m. by a scheduled cron job (chaos engineering).
- The agent is asked how to fix it and provides the command to turn off the feature flag.
- Scenario 2: API Slowdown (Background Agent):
- The agent is given a high-level directive that an API is slow.
- The agent asks itself a series of questions:
- What does the front-end service dashboard show?
- What causes high latency in traces?
- Were there any front-end service deployments?
- The agent analyzes traces and identifies the specific downstream service contributing to high latency.
- The agent uses existing knowledge about the environment (memories).
- The agent provides the fix and the commands to implement it.
- Scenario 3: On-Call Agent in Slack:
- When an alert fires, the agent responds and acknowledges the alert.
- The agent uses available tools to investigate the alert.
- The agent identifies the endpoint highlighted in the alert.
- The agent provides an initial answer with evidence, pointing to downstream services (recommendations and ad service) causing high latency.
- The agent completes the investigation and identifies that the service high latency is a direct result of severe concurrent degradation due to drum dependencies.
- The agent points to an inefficient N+1 query problem in the code.
- The agent explains what N+1 means and provides an example of how to fix it in code.
Building Multi-Agent Systems: From Simple to Complex
- Principle: Keep it simple. Not every problem requires a multi-agent system.
- Goal: Discover the right level of agentic complexity for a particular task.
- Progression:
- Single Pass LLM Call: One prompt in, one response out.
- Actor-Critic System: Adds a feedback loop with a critic LLM evaluating the actor's proposed solutions.
- LLM with Tool Use: Enables the LLM to interact with the environment using external functions (tools).
- Agentic System: Adds a planning loop and autonomous execution engine with memory and reflection.
- Multi-Agent System: Uses multiple specialized agents to address complex problems.
- Running Example: Payment system failure.
1. Single Pass LLM Call
- Process:
- Gather relevant information (logs, dashboards, incidents).
- Formulate a prompt and input it into an LLM.
- Interpret the LLM's response.
- Example:
- Paste a thousand log lines into an LLM and ask why payments are failing.
- The LLM identifies a "card expired" log message.
- Limitations:
- No internal feedback loop.
- No iterative reasoning.
- No concept of uncertainty.
- Limited to the information provided in the initial prompt.
2. Actor-Critic System
- Process:
- The actor proposes a candidate solution.
- The critic LLM examines the candidate and provides feedback.
- If the critic approves, the candidate is returned as the solution.
- If not, the actor adjusts its response based on the feedback.
- Example:
- The actor says, "Card payments are failing because the card is expired."
- The critic responds, "Card expirations are usually one-off. They wouldn't explain widespread failures."
- The actor adjusts its response to say that card expiration is a possible issue, but there may be other factors.
- Limitations:
- Still limited to the initial log lines.
- No persistent memory.
- Bound by LLM limitations (context window, math skills, hallucinations).
- No access to external knowledge.
3. LLM with Tool Use
- Definition of a Tool: Any external function that the LLM can use to interact with the environment.
- Benefits of Tools:
- Enable environment interaction.
- Provide memory.
- Enable math capabilities.
- Provide access to external knowledge.
- How Tools Look to an LLM:
- Defined in the prompt as structured interfaces.
- Each tool has a name, description, input schema, and usage protocol.
- Process:
- The LLM sees the current state, available tools, input, previous tool call results, and its own prompt.
- It generates a new candidate, possibly with tool calls.
- The execution engine executes the tool calls and appends the results.
- The LLM evaluates if the problem is solved. If not, it loops again.
- Example:
- The user says, "Payments keep failing. Check the logs."
- The LLM decides which logs to check and formulates a log query.
- The LLM sends the query to a Loki system, which returns the logs.
- The LLM analyzes the logs and identifies network issues around Stripe availability.
- Limitations:
- No state tracking.
- Struggles to manage evolving task states.
- Tool failures can be problematic.
- No real planning.
- Inefficient tool calls.
4. Agentic System
- Definition of an Agent: An LLM with tool use, a planning loop, and an autonomous execution engine.
- Key Features:
- Planning loop: The LLM can reason about a task, break it into steps, and act as its own feedback agent.
- Autonomous execution: The LLM can autonomously execute on long-running tasks.
- Memory, reflection, and internal sub-goals.
- Process:
- The agent receives a directive.
- It reasons about the directive.
- It makes tool calls (actions).
- It observes the results.
- It replans based on the results.
- It updates its memory.
- It checks if the directive is complete.
- It loops until satisfied and returns the solution.
- Example:
- The user says, "Payments keep failing."
- The agent checks dashboards to find out which service is failing.
- It queries the service logs.
- It reflects on the results and adjusts its query parameters.
- It identifies a network error connecting to Stripe.
- Limitations:
- Breadth vs. depth dilemma: One agent can't be an expert in everything.
- Planning fragility: LLMs can lose track of their main goals.
- Centralized bottleneck: Every decision flows through the main LLM.
- Error propagation: One bad guess can derail the entire conversation flow.
- Tool overload: Managing access to hundreds of tools can be overwhelming.
5. Multi-Agent System
- Rationale: Just like real-world teams have specialists, complex problems require multiple agents.
- Benefits:
- Encapsulation: Specialization of agents for specific tasks (e.g., logs, dashboards, code).
- Long-running tasks: Avoid context window limitations by distributing tasks across agents.
- Centralized decisions: Reduce the decision space for each agent.
- Errors: Add layers of protection through distributed critiquing.
- Tool context overload: Reduce the number of tools each agent needs to manage.
- Example:
- The orchestrator agent receives the message "Payments keep failing."
- It calls out to logs and dashboard agents simultaneously.
- The logs agent identifies Stripe failures due to timeouts.
- The dashboards agent identifies the checkout service as broken.
- The orchestrator agent combines the information and concludes that the checkout service is failing on Stripe timeouts.
- It calls the code agent to increase the Stripe timeout in the checkout service.
- Architectural Patterns:
- Scatter Gather: An agent kicks off subtasks, collects results, and summarizes them.
- Orchestrator Loops: An orchestrator agent calls a subset of agents, gets results, and decides on the next steps.
- Free-for-All: Agents call each other willy-nilly, sharing context and handing off tasks.
- Multi-Agent Context:
- Global Scope: Agents see the entire conversation trace.
- Advantage: No information loss.
- Disadvantage: Context window limitations.
- Local Scope: Agents only see their own input.
- Advantage: Debuggability and simplicity.
- Disadvantage: Requires the request to contain all necessary information.
- Call Stack: Agents see where they came from (the call stack).
- Compromise between global and local scope.
- Global Scope: Agents see the entire conversation trace.
- Statefulness:
- Stateful Agents: Behave like humans, remembering past interactions.
- Stateless Agents: Behave like RPCs or functions, with no memory of past requests.
- Async Activities:
- Important for UX to avoid sequential execution.
- Complex considerations for context passing and state management.
Learning and Open Questions
- Learning:
- Post-Training: Adding new training data and retraining the model.
- Pro: Higher consistency.
- Con: Expensive and difficult to maintain native smartness.
- Memory: Keeping an external scratchpad of things the agent remembers.
- Pro: Faster turnaround and more flexible.
- Con: Limited to what can be taught with memory.
- Post-Training: Adding new training data and retraining the model.
- Open Questions:
- How to tail patch a model without forgetting what it already knows.
- How to do prompting without black magic.
- Why does a prompt work on some models and not others.
- How to convince a model to introspect and know its own uncertainty.
Limitations and Future Directions
- Current Limitations:
- Reasoning abilities, especially after long chains of calls.
- Context size window.
- Long horizon tasks.
- Tools designed for humans are often terrible for LLMs.
- Human bottleneck in confirming actions.
- Future Directions:
- Moving towards agents that run constantly and solve problems on their own.
- Humans will move to higher levels of abstraction, managing agents and setting goals.
- Human tools will be replaced by agent tools.
- Emphasis on systems thinking, architecture, and fundamentals.
- Importance of alignment and good taste in agent outputs.
- Models will keep improving and enable longer horizon tasks.
- Human tools will be replaced I think by agent tools.
- We'll move from goal setting to alignment and good taste almost right and I think humans will play that role right the role of setting the goal setting the bar deciding if the output is what we need or not and you know asking the agents to do more.
AI summaries can miss context or contain errors. Check important details against the original video.