How To Prompt AI To Build the Next Billion Dollar Startup

Y CombinatorAbout 5 min readMay 30, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

Metaprompting, Prompt Engineering, Vertical AI Agents, System Prompt, Developer Prompt, User Prompt, Prompt Folding, Evals, Forward Deployed Engineer, LLM Personalities, Scoring Rubrics, Kaizen.

Parahelp's AI Customer Support Prompt

  • Main Topic: Analysis of a real-world prompt used by Parahelp, an AI customer support company powering customer support for Perplexity, Replet, and Bolt.
  • Key Points:
    • The prompt is six pages long and highly detailed.
    • It starts by defining the role of the LLM as a manager of a customer service agent.
    • The task is to approve or reject tool calls from other agents.
    • The prompt includes a step-by-step plan (steps 1-5).
    • It specifies the output format (accept/reject) for integration with other agents.
    • The prompt is structured in a markdown-like style with headings and sub-bullets.
    • It outlines how to reason about the task and provides examples.
    • Uses XML-like tags for specifying the plan, which improves LLM performance due to pre-training on similar formats.
  • Example: The prompt is used to power the AI agent responding to customer support tickets for companies like Perplexity.
  • Structure: The prompt is divided into sections like planning, creating steps, and providing examples, all formatted for clarity.

Prompt Architecture: System, Developer, and User Prompts

  • Main Topic: Emerging architecture for prompts, dividing them into system, developer, and user prompts.
  • Key Points:
    • System Prompt: Defines the high-level API of how the company operates (e.g., Parahelp's general customer support logic).
    • Developer Prompt: Adds customer-specific context and instances of the API calls (e.g., how to handle questions for Perplexity vs. Bolt).
    • User Prompt: Contains user-generated input (e.g., "generate a site with these buttons").
  • Example: Parahelp's prompt is primarily a system prompt, while a product like Replit would have a user prompt.
  • Application: This architecture helps avoid becoming a consulting company by separating general logic from customer-specific customizations.

Metaprompting and Prompt Folding

  • Main Topic: Metaprompting as a powerful tool for improving prompts, including prompt folding.
  • Key Points:
    • Prompt Folding: Dynamically generating better versions of a prompt using the LLM itself.
    • Example: A classifier prompt generates a specialized prompt based on the previous query.
    • Feeding the LLM examples where the prompt failed and asking it to improve the prompt.
  • Example: Trope helps companies like Ducky debug prompts using prompt folding.
  • Application: Useful for complex tasks where it's difficult to write precise instructions.

Examples and Test-Driven Development for LLMs

  • Main Topic: Using examples to guide LLMs, similar to test-driven development in programming.
  • Key Points:
    • Feeding the LLM hard examples that only experts could solve.
    • Adding these examples to a meta-prompt to help the LLM reason about complicated tasks.
    • This approach is particularly useful when it's hard to define precise parameters.
  • Example: Jasberry uses this approach for automatic bug finding in code.
  • Analogy: This is compared to unit testing in programming, where examples serve as tests.

LLM Escape Hatches and Debugging

  • Main Topic: Providing LLMs with escape hatches to avoid hallucinations and improve debugging.
  • Key Points:
    • LLMs tend to provide an output even when they lack sufficient information.
    • Trope's Approach: Instructing the LLM to stop and ask for more information if it's unsure.
    • YC's Approach: Adding a "debug info" parameter to the response format, allowing the LLM to report confusing or underspecified information.
  • Application: This helps developers identify and fix issues in their prompts.

Practical Metaprompting Techniques

  • Main Topic: Simple ways to get started with metaprompting for personal projects.
  • Key Points:
    • Give the LLM the role of an expert prompt engineer.
    • Ask it to critique and improve your prompt.
    • Iterate this process to refine the prompt.
    • Use larger, more powerful models (e.g., Claude 3.7, GPT-4) for metaprompting and then distill the refined prompt into a smaller model (e.g., FRO) for faster inference.
    • Use Gemini Pro's thinking traces to understand the LLM's reasoning process and identify areas for improvement.
  • Application: This can be used to quickly improve prompts for various tasks.

The Importance of Evals

  • Main Topic: Evals as the true crown jewel data asset for AI companies.
  • Key Points:
    • Evals provide context for why a prompt was written a certain way.
    • They are essential for improving prompts and understanding user needs.
    • Evals require sitting side-by-side with users to understand their workflows and reward functions.
  • Example: Parahelp considers evals more valuable than the prompts themselves.
  • Application: This is crucial for building vertical AI solutions that truly meet user needs.

The Founder as a Forward Deployed Engineer

  • Main Topic: The concept of the founder as a forward deployed engineer, drawing from Palantir's model.
  • Key Points:
    • Forward deployed engineers sit with users to understand their problems and workflows.
    • They convert these insights into clean software solutions.
    • This approach allows for rapid feedback and iteration.
    • Founders must be technical, empathetic, and great product people.
  • Example: Palantir sent engineers to sit with FBI agents to understand their investigative processes.
  • Application: This model is particularly effective for vertical AI agents, where quick iteration and deep user understanding are critical.

LLM Personalities and Scoring Rubrics

  • Main Topic: Different LLMs have different personalities and respond differently to scoring rubrics.
  • Key Points:
    • Claude is more "happy" and human-steerable.
    • Llama 4 requires more steering but can be very effective with good prompting.
    • GPT-3 is rigid and sticks closely to rubrics.
    • Gemini Pro 2.5 is more flexible and can reason through exceptions.
  • Example: Using LLMs to score investors, with different models providing different insights based on their personalities.
  • Application: Understanding these differences can help choose the right LLM for a specific task.

Synthesis/Conclusion

The video emphasizes the importance of detailed prompt engineering, the emerging architecture of system, developer, and user prompts, and the power of metaprompting for improving prompts. It highlights the critical role of evals in understanding user needs and the founder as a forward deployed engineer, deeply embedded in the user's workflow. The discussion also covers the different personalities of LLMs and how they respond to scoring rubrics. The main takeaway is that successful AI applications require a combination of technical expertise, deep user understanding, and a willingness to iterate and refine prompts based on real-world feedback. The analogy to coding in 1995 and managing a person underscores the evolving nature of prompt engineering as a discipline.

AI summaries can miss context or contain errors. Check important details against the original video.

MAKE IT YOURS

Read. Remember. Reuse.

Free tools

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.