Key Concepts
- Trustworthy AI: AI systems that are technically robust, functionally safe, human and socially appropriate, explainable and interpretable, and respectful of user privacy.
- Prompt Injection: An attack where an attacker injects data into a prompt that is interpreted as control instructions by the AI system, rather than just data.
- Agentic AI: AI systems that can perform complex tasks by breaking them down into subtasks, calling tools, and adjudicating on their completion.
- Model Context Protocol (MCP): A protocol for AI agents to interact with the outside world, which currently lacks built-in security features.
- Constitutional AI: A framework for aligning AI models with human values by using a set of principles or a "constitution."
- Retrieval Augmented Generation (RAG): A technique to reduce hallucinations by providing the AI model with a set of authoritative documents to draw information from.
- Jailbreaks: Attempts to trick an AI model into responding to questions it is not supposed to answer.
- Model Extraction: An attack where an attacker steals the functionally equivalent weights of an AI model through repeated queries.
- Hallucinations: Instances where AI models generate fabricated or incorrect information.
- Data vs. Control: A fundamental security principle where a clear distinction is maintained between data and instructions for the system.
Introduction and Background
The session features Petra Parikova interviewing Neil Daswani, Co-Academic Director of the Stanford Advanced Cybersecurity program, and Vinay Rao, CTO of ROOST and former head of the AI Safeguards team at Anthropic. Both guests have extensive experience in cybersecurity, trust, and safety roles at major tech companies. The discussion centers on the critical topic of building trustworthy AI, exploring its definition, the challenges in achieving it, and the emerging risks and defenses.
Past Interactions and Shared Experiences
Neil Daswani and Vinay Rao first met at Google, where they collaborated on combating "ad spam," specifically "click fraud." This work involved developing detection and response mechanisms in scenarios lacking clear ground truth, a challenge they note is relevant to current AI safety work where answers are not always definitively right or wrong. They co-authored a book chapter on "Online Advertising Fraud" and worked together against a bot called "Clickbot.A." This shared history highlights their long-standing expertise in dealing with nuanced and evolving security threats.
Defining Trustworthy AI
Vinay Rao outlines five key conditions for a trustworthy AI:
- Technically Robust: Reliable in delivering needed outcomes.
- Functionally Safe: Does not mislead users and provides broadly non-harmful responses.
- Human and Socially Appropriate: Avoids unwanted bias and exclusion of groups, countries, or languages.
- Explainable and Interpretable: Moves beyond "black box" perceptions to understand decision-making.
- Respectful of User Privacy: Protects data and privacy.
Neil Daswani expands on this, referencing the "TrustLLM" paper which identified eight characteristics for large language models (LLMs). He elaborates on three additional characteristics beyond Vinay's points:
- Truthfulness: Addressing "hallucinations" where models invent information. Techniques like retrieval-augmented generation (RAG) are mentioned as mitigation strategies.
- Machine Ethics: Ensuring AI responses align with generally ethical human values, a complex challenge given the difficulty in defining ethics even for humans.
- Accountability: Establishing mechanisms for responsibility when an AI model makes errors, analogous to firing an employee in a human context.
The combined list of trustworthiness characteristics includes: truthfulness, safety, fairness, robustness, privacy, machine ethics, transparency, and accountability.
Collaboration Between Security and Safety Teams
The discussion emphasizes the increasing intersection of traditional security and AI safety.
Vinay Rao explains that historically, safety teams focused on the spectrum of harms (catastrophic, high, societal, scaled), using evals and guardrails, while CISOs managed data security and supply chain attacks. However, with AI agents capable of calling external tools and accessing external data, these domains now overlap significantly. For instance, evaluating the safety of external data for internal coding tasks involves both safety (transaction-level evaluation of package safety) and security (outcome-level assurance against data exfiltration). This necessitates aligned policies, processes, and oversight mechanisms, particularly for issues like prompt injections.
Neil Daswani adds that a crucial habit for both CISOs and product/safety teams is to be "paranoid and prepared." This means operating under the assumption that systems will fail and being ready to address issues. He uses prompt injection as an example:
- Prompt Injection: An attacker injects information before or after a user's prompt, instructing the AI to ignore the original prompt and follow the attacker's commands. This can be subtle, with attackers instructing the AI to comply with their commands a small percentage of the time, making detection difficult.
- Defense Strategy: Teams must collaborate from the design stage to anticipate abuse cases, implement guardrails, and develop triage mechanisms for post-release exploitation.
Evaluation and Verification Methods for Trustworthy AI
Vinay Rao describes methods for foundation model companies:
- Post-Training: Teaching specific skills and aligning models with human values (e.g., Constitutional AI, which uses principles like the Universal Declaration of Human Rights and privacy policies).
- Alignment Training: Ensuring models adhere to human values.
- Guardrails and Classifiers: Building systems to detect harmful intent in prompts and prevent harmful model outputs, especially against prompt injections and jailbreaks.
- Monitoring Systems: Tracking user interactions and model behavior in real-world scenarios using techniques like unsupervised learning (clustering).
- Quick Response Mechanisms: Rapidly updating classifiers or adjusting system prompts based on monitoring data.
Neil Daswani highlights a key difference for AI systems: their non-deterministic nature. This requires a sample-based testing approach rather than exhaustive testing, with a focus on probabilistic measurement rather than deterministic guarantees.
Balancing Speed to Market with Security Risks
Neil Daswani uses the analogy of Facebook's evolution from "Move fast and break things" to "Move fast with stable infrastructure." He also draws a parallel to the German Autobahn, where high speeds are possible due to well-designed roads, rules, and disciplined drivers, not the absence of rules. The key is finding a balance: moving too fast without guardrails leads to problems and eventual slowdowns, while excessive focus on security can stifle innovation. The goal is to implement critical rules and guardrails that enable speed without compromising safety.
Emerging and Overhyped AI Security Risks
Vinay Rao identifies agentic AI as a significant new risk. He defines agentic AI as systems that can handle large tasks by creating subtasks, calling tools, and adjudicating outcomes. A key concern is the Model Context Protocol (MCP), which facilitates agent interaction but currently lacks security features like prompt verification or authentication of the other end. This is compared to the early days of HTTP before HTTPS, posing risks like prompt injections and data exfiltration. While MCP is valuable, it needs robust security primitives.
Neil Daswani agrees that agents present the biggest new risk, but also notes they are potentially the most overhyped. He cites a survey where 63% of security leaders identified employees giving AI agents sensitive data as their biggest cybersecurity threat. However, he suggests that basic security hygiene remains the fundamental risk. Regarding agents, he points out:
- Coding Agents: Training data for these agents may not be secure, leading to less secure code output compared to human-developed code with security reviews.
- MCP Servers: Many have been released without authentication, exposing data to attackers. Prompt injection is also a concern with LLMs behind MCP servers.
Vinay Rao further elaborates on attack classes:
- Catastrophic Risks: AI models could aid in designing highly transmissible and fatal viruses.
- Misinformation: Deepfakes and fabricated conversational elements make it hard to distinguish real from fake.
- Emotional Dependence and Mental Health: Concerns about users developing unhealthy reliance on AI.
- Societal Impact: Near-term economic impact on jobs and the potential for AI to exacerbate global divides due to its high cost, limiting access for poorer regions.
Neil Daswani focuses on prompt injection as a significant threat from a control perspective, comparing it to SQL injection in the past. He emphasizes that attackers can gain control of the AI system's output and actions. He also mentions:
- Jailbreaks: Tricking models into answering forbidden questions.
- Model Extraction: Stealing model weights through queries.
- Hallucinations: Models generating false information.
Defenses Against AI Security Risks
Prompt Injection Defenses:
- Post-Training: Reinforcement learning to establish a hierarchy of instructions (system prompt > user request > injected prompt), where lower items cannot override higher ones. However, overdoing this can make models unsteerable.
- Classifiers: Building systems to detect obviously bad intents (data exfiltration, untrusted code execution) and flag them.
- Data vs. Control Distinction: A core defense strategy.
- Anthropic's Claude Header: A header indicating content might be unverified or unsafe, intended to caution AI agents.
- Salesforce Agent Example: A vulnerability researcher found an AI agent interpreting data in a web form as instructions, highlighting the need for agents to only follow original prompts.
Other Attack Class Defenses:
- Hallucinations: Retrieval Augmented Generation (RAG), where the model is constrained to use specified authoritative documents for factual information.
- Jailbreaks: Strong system prompts and constitutional information to guide model behavior.
- Model Extraction: Referred to the AI security course for detailed defenses.
Older, "Boring" Security Practices for Modern AI
Vinay Rao emphasizes reapplying existing security knowledge:
- Cryptography: Using cryptographic signing for generated prompts to distinguish between trusted control and untrusted data, especially relevant for insecure protocols like MCP.
- Access Escalations: Applying the principle of least privilege to AI agents, ensuring subprompts have lower access than higher-level prompts to guarantee safety.
Neil Daswani adds:
- Don't Trust User Data/Input: A fundamental principle of vetting all input before significant processing.
- Security Design Reviews: Engaging security experts early in the design phase, even through informal conversations, to identify potential countermeasures and gain safety and security benefits without excessive overhead.
Conclusion and Future Outlook
Both speakers encourage the audience to ask questions and engage further. They express optimism, highlighting ongoing research in AI safety, government actions (e.g., California's AI laws, EU/UK initiatives), and the growing community of security and safety experts. They emphasize that the collective expertise and skepticism of individuals are crucial for navigating AI risks and reaping its benefits. They also call for more individuals to enter the AI safety and security fields to help address these challenges. The session concludes with a positive outlook, acknowledging both the risks and the immense potential of AI.
AI summaries can miss context or contain errors. Check important details against the original video.





