Why securing AI is harder than anyone expected and guardrails are failing | HackAPrompt CEO

Lenny's PodcastAbout 4 min readDec 29, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Fundamental Ineffectiveness of Current AI Security: Existing “guardrails” and automated red teaming provide a false sense of security and are easily bypassed.
  • Unique Challenges of AI Security: Securing AI is fundamentally different from securing traditional software due to the complexity and adaptability of AI models – “you can patch a bug, but you can’t patch a brain.”
  • Prompt Injection & Jailbreaking as Primary Threats: These attack vectors exploit vulnerabilities in LLMs and applications built on them, potentially leading to malicious outputs and data breaches.
  • The Need for Proactive, Permission-Based Security: Focus should shift from reactive defenses to proactive security measures, including granular permissioning and a deep understanding of AI vulnerabilities.
  • Escalating Risk with AI Agency: As AI systems gain more autonomy and control over real-world actions (agents, robotics), the potential for exploitation and harm increases significantly.

The State of AI Security: A Critical Assessment

The AI security landscape is currently characterized by a significant gap between perceived security and actual vulnerability. Existing defenses, primarily focused on “guardrails,” are demonstrably ineffective against prompt injection and jailbreaking attacks. This ineffectiveness stems from the infinite attack surface of LLMs and the inherent difficulty in comprehensively testing for vulnerabilities. Automated red teaming tools consistently identify weaknesses, yet human attackers consistently outperform them in bypassing defenses. The sheer number of possible attack prompts – estimated at one followed by a million zeros for a model like GPT-5 – makes exhaustive testing impossible.

Understanding the Attack Vectors: Prompt Injection vs. Jailbreaking

A crucial distinction exists between jailbreaking and prompt injection. Jailbreaking directly manipulates an LLM (like ChatGPT) to produce undesirable outputs, such as instructions for creating a bomb (Vegas Cyber Truck Bombing Plot example). Prompt injection, however, targets applications built on top of LLMs, exploiting vulnerabilities to manipulate the model through user input and override developer instructions. Examples include the Remotely.io Twitter chatbot incident, where a chatbot was tricked into posting threatening messages, and the MathGPT website compromise, which led to the exfiltration of an OpenAI API key. A recent incident involving ServiceNow Assist AI demonstrated how a chain of agent interactions could be exploited to access and potentially leak company data. The Clawed Code cyberattack showcased an AI-powered virus capable of autonomous action and API requests.

Emerging Frameworks and Concepts

Several frameworks and concepts are being developed to address AI security challenges. Camel, developed by Google, aims to restrict an AI agent’s actions based on user prompts, granting only necessary permissions (e.g., read-only access to emails). However, Camel is a framework requiring system re-architecting, not a plug-and-play solution. Constitutional AI (Anthropic’s approach) focuses on training models with a set of principles to guide behavior and prevent harmful responses, particularly regarding sensitive information categorized as SEABURN (Chemical, Biological, Radiological, Nuclear, and Explosives). Adversarial Robustness remains a challenging area, with limited progress despite ongoing research. Pdoom (Probability of Doom) represents a field within AI safety focused on assessing the risk of catastrophic outcomes. Control is a broader field dedicated to controlling potentially malicious AI, even if it actively seeks to cause harm.

The Limitations of Traditional Cybersecurity

Traditional cybersecurity approaches are insufficient for addressing the unique vulnerabilities of AI. A classical cybersecurity professional might assess an AI system (like a math problem solver) as secure based on the model itself, overlooking the potential for prompt injection to manipulate the AI into generating harmful code. This highlights the need for individuals with a strong AI security background on teams deploying AI systems. Integrating AI security with established cybersecurity principles, such as proper permissioning, is crucial, but not sufficient on its own.

The Looming Threat of Agents and Robotics & Market Correction

The risk associated with AI security is escalating as AI systems gain more agency and control over real-world actions, exemplified by AI-powered agents and robots. A predicted market correction is anticipated in the AI security industry, particularly for companies focused on ineffective “guardrails” and automated red teaming. The focus should shift towards deeper understanding of AI vulnerabilities and proactive permissioning. Open-source AI security solutions are often as effective, or more effective, than commercial offerings.

The Future of AI Security: Education and a Shift in Mindset

The conversation concludes with a call for education and a shift in mindset within the AI security community. The ultimate goal is to develop autonomous AI systems that don’t require constant human intervention, making a focus on understanding the underlying vulnerabilities of AI paramount. The speaker emphasizes that “stuff’s about to get dangerous,” highlighting the urgency of addressing these challenges.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.