Open Source Friday: Archestra – Secure platform for enterprise agents with MCP
By GitHub
Key Concepts
- MCP (Machine Communication Protocol): A protocol enabling autonomous agents to communicate with each other and external systems.
- Open Core Business Model: A business model where a company offers a core product as open-source software and sells proprietary add-ons or services.
- Lethal Trifecta: A critical vulnerability in autonomous agents arising from the combination of:
- Untrusted Context: The agent can consume data from external, potentially malicious sources.
- External Communication: The agent has the ability to send data or trigger actions in external systems.
- Sensitive Data Access: The agent can access internal, sensitive information.
- Prompt Injection: A security vulnerability where malicious input is crafted to manipulate an LLM into performing unintended actions or revealing sensitive information.
- Archestra: An open-source project building enterprise-grade software for secure MCPs, acting as a security layer for autonomous agents.
- Dual LLM Pattern (Sub-agents): A security mechanism where a separate LLM (sub-agent) sanitizes or summarizes data before it's passed to the main agent, preventing poisoning.
- Agent Technology Agnostic: Archestra can be integrated with agents built using various frameworks and technologies.
- Tool Orchestration: Archestra can manage and secure the invocation of tools (like MCP servers) by agents.
- Role-Based Access Control (RBAC): A system for managing access to resources based on user roles.
Introduction to Mavi and Archestra
Mavi, founder and maintainer of Archestra, shared his journey into open source, starting with coding at age 11 and developing a passion for open-core business models. His previous companies were acquired by Grafana Labs and Elastic, both known for their open-source enterprise software. This experience solidified his belief in building open-source solutions, stating, "I can't imagine not building something not open source anymore as it just feels natural."
Origin and Inspiration for Archestra
The spark for Archestra came from Mavi's personal experience with MCPs. While initially impressed by their capabilities, he quickly realized the significant security risks. Within minutes of connecting an MCP to platforms like Slack, colleagues could exploit prompt injections, causing the agent to act maliciously. This led to the realization that while MCP technology is powerful, it's "practically impossible to deploy in for something where like money are involved or sensitivity is involved." The core inspiration was to build an open-source solution at the infrastructure layer to enable MCP technology for "hardcore use cases with real data with real responsibility in real enterprise environment."
Core Mission and Unique Approach to MCP Security
Archestra's core mission is to provide a simpler and, crucially, a secure way to use MCPs. Mavi emphasized that without a robust security layer, MCPs are too dangerous to use with sensitive data due to the risk of hallucination and data leakage. Unlike other approaches that use LLMs to analyze agent behavior (which are still vulnerable), Archestra implements well-known security practices to prevent vulnerabilities in autonomous agents. This is delivered as a practical, plug-in open-source solution deployable within a Kubernetes cluster, compatible with agents built in any language or technology. Archestra aims to make agents safe against the "lethal trifecta."
Importance of Open Source for Archestra and Enterprises
Mavi highlighted several reasons for prioritizing open source:
- Personal Preference: "First of all we like open source. This is the honest first reason."
- Enterprise Adoption: Open source is popular and easy for enterprises to adopt. Engineers can readily access and test GitHub repositories, eliminating the need for lengthy sales cycles and POCs.
- User Proximity: Open source allows for direct interaction with end-users. Archestra's recent open-sourcing has led to immediate engagement with companies testing it, fostering rapid feedback and development.
- Trust and Verification: For enterprises building secure systems, open source allows them to "inspect it, and make sure that it is secure itself." This makes Archestra a "durable component of the supply chain" that can be trusted and verified.
The Lethal Trifecta in MCP and AI Agents
The "lethal trifecta" is a fundamental problem for autonomous agents, especially LLMs, which lack critical thinking and tend to be helpful to whoever or whatever interacts with them. The trifecta arises when an agent:
- Consumes Untrusted Data: Can ingest data from external, potentially malicious sources (e.g., public GitHub issues).
- Has External Communication Capabilities: Can send data or trigger actions in other systems (e.g., create GitHub issues, send emails).
- Accesses Sensitive Data: Can read internal, confidential information (e.g., private GitHub repositories, HR systems).
This combination creates a severe vulnerability. For instance, a GitHub MCP server, by its nature, can consume untrusted data (issues), communicate externally (create issues), and access sensitive data (private repos). Mavi stated, "even connecting one only one GitHub MCP the whole agent is getting vulnerable to the lethal trifecta and you can't go to production with that." This vulnerability is not unique to MCPs but is highlighted by them, posing a significant risk for any autonomous agent.
Secure Communication Among MCP Agents and Identity Management
Addressing secure communication between MCP agents involves technical details like using MCP itself or emerging protocols like Google's A2A. However, a more significant challenge is identity and responsibility. When agents interact, determining who is responsible for actions becomes complex, especially if one agent is launched under a user's credentials and then performs malicious actions.
Archestra's approach to this involves:
- Agent Authentication by Users: Agents should be authenticated by the users who launch them.
- Complete Traceability: A full trace of agent actions, especially when interacting with third-party systems, should be stored for investigation.
- Non-Human Identities: Agents developed by teams should operate under the identity of the group or team, not an individual. This allows for centralized control over entitlements and access. Archestra is developing an MCP orchestrator that exposes these identities to third-party systems.
Handling PII and Sensitive Data
Archestra's design fundamentally aims to prevent agents from leaking or damaging any data, including PII. While LLMs and third-party providers can help identify PII, Archestra's focus is on mitigating the entire "lethal trifecta." If the system prevents malicious data exfiltration, it inherently protects PII. The approach is to design the system to cut off the ability for agents to leak data, rather than solely focusing on PII identification within the agent's workflow.
Security Concerns with Open Source LLM Tooling and Supply Chain Attacks
The discussion touched upon the risks associated with open-source LLM tooling, including:
- Malicious Prompts in MCP Servers: Many MCP servers are supported by few, often anonymous, maintainers. These servers can contain prompt injections that, when an agent uses them, can compromise the entire agent. Pinning to the latest version of such MCPs can introduce vulnerabilities overnight.
- Running MCP Servers Locally: Running MCP server binaries on personal laptops is extremely dangerous and can lead to system compromise.
- Training on Malicious Data: Archestra's approach, using static guardrails and its proxy mechanism, aims to mitigate risks even if models are trained on malicious data. The system's safeguards prevent malicious tool calls, regardless of the model's training data.
Archestra Features and Protection Against the Lethal Trifecta
Archestra is designed as a pluggable component that sits between the LLM and the agent, acting as a proxy. Its key features for security include:
- Agent Technology Agnostic Deployment: Can be deployed as a Docker container or Helm chart alongside any agent, requiring only a change in the LLM endpoint URL (e.g., from OpenAI to Archestra's proxy).
- Tool Interception and Policy Enforcement: Archestra intercepts tool calls made by the agent. It parses available tools and allows administrators to define policies specifying which tools are allowed in "untrusted contexts" (i.e., when the agent might be compromised).
- Policy Configuration: Policies can be defined by platform teams or agent development teams. Archestra supports RBAC for policy management and is developing an automatic policy builder.
- Cataloging MCP Servers and Tools: Archestra maintains a catalog of known MCP servers and their tools, enabling pre-defined security policies for popular integrations.
- Granular Policy Control: Policies can be highly granular, distinguishing between trusted and untrusted contexts for specific tools or even based on the source of data (e.g., emails from colleagues vs. external domains).
- Dual LLM Pattern (Sub-agents): This feature allows agents to interact with potentially dangerous or sensitive data safely. A separate LLM (sub-agent) sanitizes or summarizes the data, answering specific questions from the main agent without exposing the raw, potentially poisoned content. This preserves agent functionality while mitigating risk.
Live Demo Example: GitHub MCP and Lethal Trifecta Prevention
During the demo, a simple NA10 agent with a GitHub MCP server was shown to be vulnerable. A malicious issue containing a prompt injection instructed the agent to read from a private repository and create a new issue in a public repository with the summarized content. This successfully exfiltrated data.
When Archestra was introduced as a proxy, the agent's LLM endpoint was switched to Archestra. Archestra identified the GitHub tools and allowed configuration of policies. By disallowing the create_issue tool in untrusted contexts, Archestra blocked the malicious action.
The demo then showcased the Dual LLM pattern. The agent was instructed to read the malicious issue and simultaneously create a new issue. Archestra's sub-agent quarantined the content of the malicious issue, questioned the quarantine about its general nature (e.g., "feature requests"), and only then allowed the main agent to proceed with creating the new issue. This prevented the main agent from being poisoned while still enabling it to perform its intended action.
Performance and Latency
Archestra adds approximately 25-50 milliseconds to streaming latency without the Dual LLM pattern, which is "almost invisible" to the user. The Dual LLM pattern can add 2-3 seconds due to the session of questions and answers between LLMs. This is considered a worthwhile trade-off for preventing security breaches. Optimization is ongoing, including using faster, smaller models for sub-agent tasks.
Future Vision and Community Engagement
Archestra's vision is to be a production-ready, enterprise-grade open-source security solution for MCPs. Key future developments include:
- Enhanced Security Mechanics: Expanding security patterns beyond the current offerings.
- Production Readiness: Focusing on robust logging, versioning, and stability for enterprise deployment.
- MCP Orchestrator: A significant upcoming feature that will manage authentication, identities, and the execution of MCP servers at the network layer, eliminating the need for agent code modifications. This includes tackling the complex authentication challenges within MCP protocols. This orchestrator is slated for announcement at CubeCon in Atlanta.
- Community Involvement: Archestra is highly community-driven. They encourage joining their Slack channel, visiting their website (archestra.ai), and contributing via GitHub issues. The entire engineering team is active in the Slack community.
Conclusion
Archestra is addressing a critical security gap in the rapidly evolving landscape of autonomous agents and MCPs. By providing an open-source, pluggable security layer, it enables enterprises to leverage the power of these technologies without succumbing to the inherent risks of the "lethal trifecta." The project's focus on practical implementation, robust security features like policy enforcement and the Dual LLM pattern, and a strong commitment to community engagement positions it as a significant player in securing the future of AI agents.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

Why Does This Guy Appear In Kids Videos?
sphynx

TIC en las Organizaciones - Electiva Complementaria II Unisimon
Julieth Güell S

How to Tame Your Advice Monster | Michael Bungay Stanier | TED
TED

Margaret Heffernan: Why it's time to forget the pecking order at work
TED

The importance of psychological safety: Amy Edmondson
The King's Fund

What Is Psychological Safety?
Harvard Business Review

13-Conflict Management: Listening in Conflict
Deliberate Development