The two ways people trick AI

Lenny's PodcastAbout 3 min readDec 25, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Jailbreaking: Directly manipulating a Large Language Model (LLM) through crafted prompts to bypass safety protocols and generate harmful outputs.
  • Prompt Injection: Exploiting vulnerabilities in applications using LLMs by crafting prompts that override the developer’s intended instructions and manipulate the model’s behavior.
  • Large Language Model (LLM): A type of artificial intelligence that uses deep learning algorithms to understand and generate human language.
  • Developer Prompt: The initial instructions given to the LLM by the application developer, defining its role and limitations.

Defining Jailbreaking and Prompt Injection: A Comparative Analysis

The core distinction between jailbreaking and prompt injection lies in the number of actors involved and the attack vector. Jailbreaking represents a direct interaction between a malicious user and the underlying Large Language Model (LLM), such as ChatGPT. The attack relies solely on crafting a sufficiently complex or deceptive prompt to circumvent the model’s built-in safety mechanisms. An example provided is a user submitting an exceptionally long and malicious prompt to ChatGPT, successfully tricking it into generating harmful content – specifically, instructions for constructing an explosive device. This highlights the vulnerability of LLMs to cleverly designed inputs that exploit weaknesses in their training data or algorithmic safeguards.

The Mechanics of Prompt Injection: Exploiting Application Layer Vulnerabilities

Prompt injection, conversely, is a more nuanced attack that targets applications built on top of LLMs. The scenario presented involves a website, “write a story.ai,” designed to generate stories based on user-provided ideas. A malicious user doesn’t directly attack the LLM itself, but rather exploits the application’s interface. They achieve this by crafting a prompt that instructs the LLM to disregard the developer’s original instructions (the “developer prompt”) and instead fulfill a malicious request – again, exemplified by requesting bomb-making instructions.

This demonstrates that prompt injection isn’t about breaking the LLM’s core functionality, but about hijacking its behavior within a specific application context. The developer prompt acts as a set of constraints and guidelines for the LLM, and prompt injection aims to override these constraints.

Key Differences Summarized

The speaker explicitly states the key difference: “In jailbreaking, it’s just a malicious user and a model. In prompt injection, it’s a malicious user, a model, and some developer prompt that the malicious user is trying to get the model to ignore.” This clarifies that jailbreaking is a direct model attack, while prompt injection is an application-level attack leveraging the LLM.

Implications and Actionable Insights

The discussion underscores the importance of robust security measures at both the LLM level (to mitigate jailbreaking) and the application level (to prevent prompt injection). Developers utilizing LLMs must carefully consider the potential for malicious prompt manipulation and implement safeguards to protect against unintended or harmful outputs. This includes careful prompt engineering, input validation, and potentially, techniques to isolate the LLM from untrusted user input.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.