How METR measures Long Tasks and Experienced Open Source Dev Productivity - Joel Becker, METR

AI EngineerAbout 4 min readJan 19, 2026Watch original
THE SUMMARYAI-generated

Key Concepts

  • Compute Growth & AI Capabilities: A direct relationship exists between compute growth and the time horizons of AI tasks; slowing compute growth will decelerate AI progress.
  • AI Safety & Monitoring: Limiting AI task duration based on monitoring levels is a potential safety mechanism.
  • AI Limitations & Real-World Application: Current AI systems struggle with real-world complexity due to a mismatch between their cognitive architecture and the human-designed environment.
  • Automating AI Research: Fully automating AI research requires advancements in both software and hardware (chip design and production).
  • Capability Measurement Challenges: Current AI benchmarks have limitations and require triangulation with other evaluation methods.

Forecasting AI Progress & Compute Constraints (Part 1)

The discussion begins with the premise that AI capability advancement is directly proportional to compute growth. If compute growth halves, the time it takes AI to complete tasks (time horizon) will also halve. While acknowledging the complexity of forecasting compute growth, the speaker believes a slowdown is likely due to physical (power limitations, as noted in an EPO report) and, more importantly, financial constraints on large organizations. This proportionality holds until a “software-only singularity” is achieved, a point where software advancements drastically reduce the need for raw compute. Log-linear plots are presented as a historically accurate forecasting tool for AI, and the default expectation should be for these trends to continue unless disrupted.

A key challenge identified is that as time horizons shrink, tasks become too short to meaningfully test AI capabilities. AI development is primarily a human-machine iteration loop, with reliability becoming more crucial than simply reducing task completion time. The hypothetical milestone of “one-work” illustrates the potential impact of delayed capabilities. Studies were cited, including an open-source developer study (16 developers in a hackathon) showing a small, statistically insignificant speedup with AI assistance (approximately 4 percentile points), and a Meta presentation highlighting a J-curve in developer productivity with AI agents, requiring 3-6 months for noticeable improvement. The HA Haskell compiler example demonstrates a high quality bar and rigorous review process in some open-source projects, potentially limiting the immediate impact of AI assistance. LinkedIn’s complex data infrastructure (5,000 “impressions” tables) highlights a potential area for AI value, hampered by data quality and complexity.

Existential Risk, Capabilities & the “Neurodivergent” AI (Part 2)

The discussion shifts to the feasibility of controlling AI as capabilities double, with short-term solutions (1-2 years) appearing manageable, but beyond 3 years posing a significant challenge. Current AI, specifically GPT-5, is not considered an existential threat due to its inability to perform complex data science on ambiguous datasets; a 90% success rate on challenging tasks is more concerning. A proposed safety mechanism involves monitoring AI activity, allowing longer tasks only without monitoring and shorter tasks under surveillance, creating a variable time horizon.

Current capabilities measurement methods, like algorithmic scoring on MMLU, are critiqued, emphasizing the need to assess whether work can be built upon, not just whether a task is completed. A central argument is that current AI systems are akin to “neurodivergent individuals” – exceptionally skilled in specific areas but struggling with the complexities of the real world, designed for human capabilities. This explains their difficulty automating tasks requiring common sense or flexible problem-solving, as demonstrated by consistent failures in platforms like Agent Village AI. The discussion extends to automating robotics and chip production, questioning its necessity for a self-sustaining AI economy and its feasibility (ranging from decades to centuries). The HA Haskell compiler example is revisited, highlighting the exceptionally rigorous code review process where the median pull request requires zero minutes of post-review work.

Conclusion

The analysis presented suggests a cautious outlook on the continued rapid advancement of AI. While acknowledging the impressive capabilities of current models like GPT-5, the discussion emphasizes the critical role of compute growth, the limitations of current AI systems in real-world applications, and the significant challenges in automating AI research and development. The proposed safety mechanisms and the analogy of AI as “neurodivergent” highlight the need for a nuanced understanding of AI’s strengths and weaknesses, and a focus on robust monitoring and realistic expectations for future progress. The core takeaway is that continued AI advancement is not guaranteed and will likely be constrained by factors beyond simply increasing model size and compute power.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.