Devin 2.0 and the Future of SWE - Scott Wu, Cognition

AI EngineerAbout 6 min readJul 26, 2025Watch original
THE SUMMARYAI-generated

Key Concepts:

  • AI Agents: Software programs designed to perform tasks autonomously.
  • Moore's Law for AI Agents: The observation that the capability of AI agents, measured by the amount of uninterrupted work they can perform, is doubling at an accelerating rate.
  • Repetitive Migrations: Code refactoring tasks that involve applying a consistent set of steps across a large codebase (e.g., JavaScript to TypeScript).
  • Instruction Following: The ability of an AI agent to accurately execute a predefined sequence of steps.
  • Knowledge/Memory: The capacity of an AI agent to retain and apply learnings from previous tasks to improve future performance.
  • Codebase Intelligence: An AI agent's understanding of the structure, relationships, and context within a software project.
  • Iterative Workflow: A collaborative process where humans and AI agents work together, with humans providing guidance and feedback at key points.
  • Asynchronous Testing: The ability of an AI agent to independently test its code changes and validate their correctness.

1. Moore's Law for AI Agents

  • The speaker introduces the concept of "Moore's Law for AI Agents," defining it as the rate at which an AI's capability to perform uninterrupted work increases.
  • He contrasts the progress of general language models (like GPT-3, GPT-3.5, and GPT-4) with that of AI agents specifically designed for coding tasks.
  • While general language models see a doubling of capability roughly every seven months, AI coding agents are experiencing a doubling every 70 days (approximately two to three months).
  • This translates to a potential 16x to 64x increase in the amount of coding work an AI agent can handle in a single year.

2. Evolution of AI Agent Capabilities: Tiers of Development

  • The speaker outlines the evolution of AI agent capabilities over the past year, dividing it into distinct tiers, each marked by a significant advancement in what AI agents can reliably accomplish.
  • He emphasizes that the "future interface" and "most important capabilities" for AI agents are constantly changing, evolving every two to three months as new tiers are reached.

3. Tier 1: Repetitive Migrations (Summer of Last Year)

  • The initial major use case for AI agents was in performing repetitive code migrations, such as converting JavaScript to TypeScript or upgrading framework versions (e.g., Angular, Java).
  • These tasks are characterized by a clear set of steps that need to be applied consistently across a large number of files.
  • The primary challenge at this stage was instruction following: ensuring the AI agent could reliably execute the predefined steps.
  • The solution involved building systems like "playbooks," which allowed users to outline a clear sequence of actions for the agent to follow.
  • Another key development was the implementation of knowledge/memory, enabling the agent to learn from feedback and improve its performance on subsequent migrations.

4. Tier 2: Isolated Bug Fixes and Features (End of Summer/Fall)

  • As AI agents became more capable, they could handle more general bug fixes and feature implementations, albeit still relatively isolated in scope.
  • An example given is asking the agent to reorder items in a select dropdown.
  • The key requirement here was the ability to set up and work with the repository, including running linters and CI checks.
  • The solution involved creating a system for taking snapshots of the repository, allowing for easy setup, rollback, and reloading.

5. Tier 3: Broader Bugs and Requests (Fall)

  • This tier saw AI agents tackling more complex tasks that required understanding and modifying multiple files across the codebase.
  • These tasks often involved diagnosing issues, understanding code relationships, and making consistent changes across different parts of the project.
  • A crucial development was the ability to analyze code not just as text, but as a hierarchical structure, leveraging call hierarchies, language servers, and Git commit history.
  • This allowed the agent to understand the context of the code and make more informed decisions.
  • Integration with communication platforms like Slack became important, enabling users to easily assign tasks to the AI agent.

6. Tier 4: Complex Tasks and Iterative Workflows (Spring of This Year)

  • As tasks became more complex, it became necessary for humans to collaborate more closely with AI agents, providing guidance and feedback throughout the process.
  • The speaker notes that for harder tasks, the human doesn't necessarily know everything that needs to be done at the start.
  • This led to the development of tools like Deep Wiki and advanced search capabilities, allowing users to explore the codebase with the agent and understand the task requirements before execution.
  • The workflow shifted towards a more iterative model, with humans monitoring the agent's progress and providing input at key points.
  • The release of Devon 2.0 and the in-IDE experience reflected this shift towards a more collaborative approach.

7. Tier 5: Backlog Killing and Autonomous Task Execution (June)

  • The latest tier focuses on enabling AI agents to tackle entire backlogs of tasks autonomously.
  • This requires the agent to scope out tasks, understand requirements, decide when to seek human input, and work across multiple repositories.
  • A key factor is the agent's confidence in its understanding of the task, allowing it to proceed autonomously when appropriate.
  • Testing becomes crucial at this stage, with the agent needing to independently test its code changes and validate their correctness.
  • Asynchronous testing and iterative feedback loops are essential for ensuring the quality of the delivered code.

8. The Future: Projects and Beyond

  • The speaker concludes by discussing the future direction of AI agents, focusing on their ability to tackle entire projects and even larger, more strategic initiatives.
  • He reiterates that the specific challenges and requirements change with each doubling of capability, requiring continuous innovation in tooling and methodologies.
  • He emphasizes that each "2x" improvement is different, requiring advancements in areas like human collaboration, feedback mechanisms, debugging capabilities, and long-term decision-making.

9. Notable Quotes:

  • "Moore's law for AI agents... the capability or the capacity of an AI [is measured] by how much work it can do uninterrupted until you have to come in and step in and intervene or steer it."
  • "Every time you get to the next tier, the bottleneck that you're running into or the most important capability or the right way you should be interfacing with it, like all these actually change at each point."
  • "AI has always done these more boilerplate tasks and the more tedious stuff, the more repetitive stuff, and we get to do the the the more fun creative stuff."

10. Technical Terms and Concepts:

  • PMF (Product-Market Fit): The degree to which a product satisfies market demand.
  • GA (Generally Available): A software release that is available to the general public.
  • L2 Experience: A more interactive and collaborative experience, often involving direct communication and feedback.
  • CI (Continuous Integration): A software development practice where code changes are frequently integrated and tested.
  • Linter: A tool that analyzes code for potential errors, style violations, and other issues.
  • Diff: A comparison of two versions of a file, highlighting the changes made.

11. Synthesis/Conclusion:

The speaker provides a compelling overview of the rapid advancements in AI coding agents, highlighting the exponential growth in their capabilities and the evolving challenges and solutions at each stage of development. He emphasizes the importance of adapting to the changing landscape and focusing on the most pressing bottlenecks to unlock the full potential of AI in software engineering. The key takeaway is that AI agents are rapidly transforming the software development process, automating increasingly complex tasks and enabling developers to focus on higher-level, more creative work.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.