Prompt Learning and Evaluation for Coding Agents
Key Concepts:
- System Prompt Learning: Iteratively refining system prompts for Large Language Models (LLMs) based on feedback, analogous to human learning.
- LLM as a Judge (Eval): Utilizing an LLM to evaluate code generated by another LLM, providing detailed explanations of failures.
- Meta Prompt: A prompt that takes the original system prompt, rules, input, LLM evaluation, and explanation as input to generate an improved system prompt.
- SWEBench: A benchmark suite for evaluating code generation models, consisting of 150 examples in this case.
- GEA (or JEA): A prompt optimizer from DSPy that uses English feedback to improve prompts.
- Sample Efficiency: The amount of data required to achieve a desired level of performance.
I. The Importance of System Prompts
The speaker begins by highlighting the often-overlooked importance of system prompts in the performance of frontier coding models like Claude and Gemini. A viral tweet showcasing the extensive length of Claude’s system prompt illustrates that these prompts are not static but are continuously iterated upon. Karpathi’s concept of “system prompt learning” is introduced, drawing a parallel to the movie Memento, where continuous feedback and re-learning are essential due to memory loss. This iterative process is presented as a core component of successful coding agents.
II. Prompt Learning vs. Reinforcement Learning (RL)
A comparison is drawn between prompt learning and Reinforcement Learning (RL) to illustrate the advantages of the former for building coding agents.
- RL Analogy: A student taking an exam receives only a scalar reward (e.g., 70%, 80%). They must independently deduce how to improve. This process is described as potentially sample inefficient, time intensive, and data hungry, requiring a dedicated data science team.
- Prompt Learning Analogy: The student receives not only a score but also detailed English feedback explaining strengths and weaknesses, specific concepts missed, and areas for further study. This allows for more targeted improvement.
The speaker argues that prompt learning may be a more suitable paradigm for teams building agents, given the inherent capabilities of LLMs.
III. Methodology: Implementing Prompt Learning with Claude and Gemini
The methodology involved iteratively improving the system prompts of both Claude and Gemini (specifically Claude Code and Gemini Code) using a process centered around evaluation and feedback.
- Initial Benchmarking: Both models were initially benchmarked on SWEBench (150 examples) with no modifications to their system prompts. Claude Code achieved approximately 40% resolution of GitHub issues, while Gemini Code achieved around 30%.
- Code Generation & Unit Testing: The coding agents were tasked with solving problems from SWEBench, generating code patches. These patches were then run through unit tests.
- LLM as a Judge (Eval): The results of the unit tests (pass/fail) were fed into an LLM configured as a judge. This LLM was prompted to provide a detailed explanation of why the code failed, identifying specific error categories (e.g., parsing errors, handling specific libraries). The speaker emphasizes the importance of well-crafted “eval prompts” for generating insightful explanations.
- Meta Prompt & System Prompt Iteration: The original system prompt, existing rules (initially empty for both models), the problem statement, the LLM’s evaluation, and the explanation of the failure were all combined into a “meta prompt.” This meta prompt was used to generate an updated system prompt with new rules based on the identified errors.
- Re-Benchmarking: The process was repeated, re-benchmarking the models on SWEBench with the updated system prompts.
IV. Results and Performance Improvements
The prompt learning process yielded significant improvements in performance:
- Claude Code: Improved by 5% in GitHub issue resolution.
- Gemini Code: Improved by 15% in GitHub issue resolution.
These improvements were achieved using only 150 examples of training data. The speaker emphasizes the potential impact of this approach for improving agent performance. The tests were also run on BBH and other software engineering datasets, though specific results weren’t detailed.
V. Prompt Learning vs. GEA (DSPy)
A comparison was made between the speaker’s prompt learning approach and GEA (or JEA), a prompt optimizer from DSPy. While both utilize English feedback, the speaker’s method required significantly fewer iterations and rollouts. The key differentiator was the emphasis on developing high-quality “eval prompts” that generated detailed and informative explanations of failures.
Quote: “...eval really, this was super critical for us to be able to get this to work.”
VI. Technical Details & Considerations
- SWEBench: A benchmark suite consisting of 150 software engineering problems.
- Cloud MD & Client Rules: Mechanisms for appending rules to the system prompts of Claude Code and Gemini Code, respectively.
- LLM as a Judge Eval Prompt: The prompt used to configure the LLM to evaluate code and provide explanations. The quality of this prompt is crucial.
- Meta Prompt: The prompt used to generate updated system prompts based on feedback.
VII. Conclusion
The speaker concludes that prompt learning, particularly when combined with a robust evaluation framework (LLM as a Judge), offers a powerful and efficient method for improving the performance of coding agents. The approach is presented as a viable alternative to more complex methods like RL, especially for teams without extensive data science resources. The emphasis on iterative refinement based on detailed feedback, coupled with well-designed evaluation prompts, is highlighted as the key to success. The speaker encourages further exploration of these techniques and invites attendees to learn more through their blog and consider joining their team.
AI summaries can miss context or contain errors. Check important details against the original video.