METR

About
METR (pronounced ‘meter') is a nonprofit that evaluates frontier AI models to help companies and wider society understand AI capabilities and what risks they pose. To better understand what we do, take a look at these examples of our research.
- Measuring AI Ability to Complete Long Tasks
- Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity
- MALT: A Dataset of Natural and Prompted Behaviors That Threaten Eval Integrity
- CoT May Be Highly Informative Despite “Unfaithfulness”



A Technical Challenge for the 3blue1brown audience
Optimize a GPU kernel—or try to "reward hack" your way to a high score. Just don't get caught by the AI reviewer.
The Kernel Optimization ChallengeFeatured work
Uplift Study
We conducted a randomized controlled trial (RCT) measuring how early-2025 AI tools affected the productivity of 16 experienced open-source developers working on large, mature codebases (avg. 5 years xp w/ repo). Developers completed 246 real issues, which were randomly assigned to either allow or disallow AI usage. Surprisingly, we found that when developers used AI tools, they took 19% longer—AI slowed them down. This contrasted sharply with both the developers' own expectations and expert forecasts (24% and 38-39% shorter time to complete tasks when allowed to use AI, respectively).

Read the full result
Time Horizon Study
We propose measuring AI performance in terms of the length of tasks AI agents can complete. We show that this metric has been consistently exponentially increasing over the past 6 years, with a doubling time of around 7 months. Extrapolating this trend predicts that, in under a decade, we will see AI agents that can independently complete a large fraction of software tasks that currently take humans days or weeks.

Read the full result
Message from Grant
A lot of people in AI circles will recognize METR's name from a few of their now-famous studies. They are the ones who introduced the idea of a 7-month doubling time for the duration of tasks that models can reliably perform, and in fact they were the ones to put focus on that as a metric in the first place. They are also the ones who put out the early 2025 study on whether AI tools actually improved the productivity of experienced engineers.
In meeting their team and learning about their culture, one person summed it up as a mix between a PhD program and a startup, which absolutely tracked with the energy I was getting. Another said the team energy was essentially that of a math camp. But this is just the social vibe; the substance comes from a shared mission, with everyone united by a deep care for AI safety, combined with a shared principle of directing that concern toward research that is real, grounded, and useful.