How we hacked YC Spring 2025 batch’s AI agents — Rene Brandel, Casco

AI EngineerAbout 4 min readJul 31, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

AI Agents, Red Teaming, Agent Security, Crosser Data Access, IDOR (Insecure Direct Object Reference), Authentication, Authorization, Code Execution Sandboxes, Server-Side Request Forgery (SSRF), System Prompt, Tool Definitions, Service Endpoint Discovery, Metadata Discovery, Firecracker, Policy Puppeteering.

Agent Security: Hacking AI Agents

Introduction

Renee, CEO of Casco, a YC company specializing in red teaming AI agents and apps, discusses the importance of agent security beyond just LLM security. She shares experiences from her past, including building voice-to-code systems 10 years ago, and highlights the evolution of AI agent architecture.

Evolution of AI Agent Architecture

  • Past: Complex, involving multiple cloud providers (e.g., IBM Watson, Microsoft Lewis) and piecing together various services.
  • Present: More normalized stack with a server front end, API server, LLM, tools, and data sources. This normalization simplifies development but necessitates a focus on security.

The Need for Agent Security

While prompt injection and harmful content generation are important concerns, real damage often occurs due to vulnerabilities across the entire system. Agent security encompasses all potential attack vectors.

Hacking AI Agents: A Case Study

Casco hacked 7 out of 16 live agents within 30 minutes each, revealing three common security issues. This was done to create a splashy launch at Y Combinator.

Crosser Data Access (IDOR)

  • Problem: Agents with tools to look up user info by ID or document by ID are vulnerable to IDOR attacks. Attackers can guess or find IDs (e.g., in product demo video URLs) and access sensitive information.
  • Example: Leaking a company's system prompt revealed tools for looking up user info and documents by ID. User IDs found in a demo video URL were used to access personal information (email, nickname). Interconnected IDs (user, chat, document) allowed traversal of the entire system.
  • Fix: Implement both authentication (validating the token) and authorization (access control matrix) to ensure the requestor has the necessary permissions. Role-level security is crucial.

Agents as Users, Not API Servers

  • Key Idea: Agents should be treated as users, not API servers, in terms of security.
  • Implications:
    • LLMs should not determine authorization patterns.
    • Agents should not act with service-level permissions.
    • Inputs and outputs should be sanitized.

Code Execution Sandboxes

  • Problem: Agents with code tools can be exploited if the code execution sandbox is not properly secured.
  • Example: An agent allowed writing Python files and reading files. Attackers bypassed security checks by:
    1. Reading the file system to find the app.py file.
    2. Overwriting the app.py file (which contained security checks) with empty strings.
    3. Gaining arbitrary code execution, leading to Bitcoin mining and access to customer data via BigQuery.
  • Lateral Movement: Arbitrary code execution allows attackers to discover other devices and resources on the network (service endpoint discovery, metadata discovery) and fetch service tokens.
  • Solution: Use out-of-the-box code sandbox solutions (e.g., etb) with observability and easy integration. Avoid rolling your own. Firecracker is mentioned as a good isolation layer. Containers alone are not sufficient for isolation.

Server-Side Request Forgery (SSRF)

  • Problem: Agents can be tricked into calling unintended endpoints, revealing sensitive information.
  • Example: An agent that creates databases pulled the database schema from a private GitHub repository. Attackers injected a malicious Git repo URL, causing the agent to send Git credentials to the attacker's server. This allowed the attacker to download the entire codebase from the private repository.
  • Fix: Sanitize inputs and outputs to prevent malicious URLs or other exploits.

Key Takeaways

  1. Agent security is bigger than just LLM security: Consider all potential threat vectors across the entire system.
  2. Treat agents as users: Apply user-level security practices, including authentication, authorization, and input/output sanitization.
  3. Don't roll your own code sandbox: Use established, secure solutions to prevent arbitrary code execution vulnerabilities.

Additional Resources

  • Casco: Offers AI agent red teaming services.
  • Hidden Layer: Has a blog post on policy puppeteering attacks.

Conclusion

Securing AI agents requires a comprehensive approach that considers all potential vulnerabilities, treats agents as users, and leverages established security solutions. By addressing these key areas, developers can build more secure and reliable AI agent systems.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.