26 Comments
User's avatar
Keller Scholl's avatar

You're far too harsh:

1. On a fixed budget, I think a higher number of tasks, even with a lower number of participant-completions, was the right move. Yes, they've got very few, but 1 would be non-zero informative and 3 is in fact something. And I'd much rather 3 people doing 100 tasks than 100 people doing 3 tasks.

2. 50% reliability seems reasonable if your plan is "set one of my 5-8 simultaneously running agents off to do something, check back with it in a few hours". And that seems to be the workflow of many AI users I know. So long as the task is checkable...

3. One of the benefits of recruiting from people in your social network, rather than randoms you're paying, is that they're not going to deliberately slow their work, and you have some evidence of competence.

Nathan Witkin's avatar

1. Sure, I'm not claiming that larger sample / fewer-tasks would've been better. But again, this sample is far too small to be informative of what baseline times might be within a large, random sample of software engineers, so this doesn't really engage with my main critique. I'm not sure I buy that it was impossible for them to do better sampling within their budget, but even if it wasn't, I think the proper response would either have been a) to use some other within-budget method, or b) to use the method they did, but to be substantially more upfront about the limited generalizability of their results.

2. Yes and no. I question whether this approach, at 50% reliability, would result in a net-productivity increase for the average coder, esp. when you consider you're burning money on tokens the whole way. But regardless, my understanding is that METR is much more interested in forecasting high-risk scenarios in which AI substitutes for a large share of software engineering labor. In that context, the 80% graphs are more relevant.

3. I take the point that this probably makes slow-rolling less likely, but I doubt it eliminates it altogether. I don't buy that, just because you know some people at METR (possibly through a couple degrees of separation) you're not going to work just a bit more slowly if you're being paid to do so. But again I think you're missing the stronger critique. METR also found that their baseliners were ~5-18x slower than their repository maintainers (although this was on tasks the latter were uniquely well equipped for), so there is other strong evidence baseliners were slower than would be realistic for a software engineer doing their normal job.

Keller Scholl's avatar

3. One of the reports I hear over and over again from others, and experience myself, is that models are useful in domains outside of one's core expertise (but where validation is still possible). I don't think "repository maintainer in their own repository" is necessarily the standard, when so much of the work that programmers do involves grasping a new domain and solving problems in it.

Elias Schmied's avatar

Useful, thanks (although a bit harsh for my liking).

"I expect some of you will want to say something like: “well, even if the specific times are off, the METR graph still shows model capabilities doubling at a frightening rate.” No. Do not try this. If the specific times are wrong, so are their distribution. That is to say: were we to collect baselines from a much larger, more representative sample, some tasks might fall into different completion-time buckets. For instance, tasks METR found took 2-3 hours might turn out, in this larger sample, to take 1-1.5 hrs. But, unless all completion times shift downward by the exact same proportion, this would compress the overall distribution of task completion times. That would in turn change the rate at which A.I. models moved up METR’s y-axis."

I'd have liked to see more elaboration on this point, since it's the most important one in the whole post. I'm not sure I have a good grasp of how strong this argument is.

Edit: ah, your exchange with Jitse in the comments does some of this.

Elias Schmied's avatar

"METR’s own research suggests that their sample underrepresents domain experts relative to engineers that are experienced, but working outside of their speciality. If a larger, random sample corrected for this imbalance, that could have the effect of compressing METR’s distribution of task completion times, since, while very short tasks presumably have more stable floors, “long” tasks may turn out much shorter when tackled by engineers with pertinent expertise. This would in turn make the AI capability growth curve much shallower, lengthening what METR refers to as “doubling times” — how long it takes for the task completion times at which models achieve 50% success rates to double."

Okay, this is convincing. Amazing, thanks.

Anatol Wegner, PhD's avatar

If anyone is interested I did a critical analysis of the METR paper from a statistics point of view a while ago: https://aichats.substack.com/p/are-ai-time-horizon-doubling-every?r=4tn68o

Charlie Harrison's avatar

Thank you, Nathan, for the interesting post

Jitse Goutbeek's avatar

I think you make very convincing arguments that the numbers of the y-axis might be inflated, compared to real world use we might expect. However, none of your points diminish anything about the exponential slope which is the headline results. This slope is also visable in the 80% completion graph and we at least see clear increases in the messy task benchmark.

The crucial part for measuring the slope is not exactly how they measured task-time completion but that this is done consistently across models they measure, this seems to have been the case. All of your arguments apply just as much to measuring the capabilities of GPT-2 as GPT-5. If we accept that the slope is exponential with pretty fast doubling time that means that at the very least the models are very quickly getting better at something.

Based on your valid critiques you can disagree about how good the models are at automated tasks right now compared to humans, but even if these are inflated by a hundred times (which seems like a lot) that would only lengthen timelines by 3-4 years based on doubling. On a human life that is not a lot.

In some ways it is a big deal if the Y-axis is wrong, the further we are away the more time we have and the longer the trend would need to continue for it to be a big deal, creating additional uncertainty. However in many of the most important ways the slope matters much more than the level. People working with these models to code tend to have some experience with how good the models are right now, and if we are indeed looking at continued exponential improvement in task lengths it is easy to see how this could help automate AI-research itself at some point in the relatively near future. This is a huge deal whether these futures are in 5 or 10 or even 20 years.

Sometimes AI-timeline debates assume such wild futures that the idea of fundamental transformation of the economy in 10-20 years seems timid. Stepping outside of these debates we can see this would clearly still be an extremely big deal and worth discussing, including with this graph which is mainly used to illustrate the exponential (not really the details of the Y-axis).

I wonder if you also believe the slope to be wrong, and if so why? None of your points seem to address this.

Nathan Witkin's avatar

"None of your points seem to address this."

There is a lengthy paragraph aimed at preempting exactly this comment. Pasting it below for your convenience.

I expect some of you will want to say something like: “well, even if the specific times are off, the METR graph still shows model capabilities doubling at a frightening rate.” No. Do not try this. If the specific times are wrong, so are their distribution. That is to say: were we to collect baselines from a much larger, more representative sample, some tasks might fall into different completion-time buckets. For instance, tasks METR found took 2-3 hours might turn out, in this larger sample, to take 1-1.5 hrs. But, unless all completion times shift downward by the exact same proportion, this would compress the overall distribution of task completion times. That would in turn change the rate at which A.I. models moved up METR’s y-axis. We have no way of ruling out this or similar possibilities, meaning that—given METR’s tiny sample—we cannot use their results make general inferences as to how much better models are getting, or how fast. This is why I call the benchmark (nearly) useless: it is being sold as licensing inferences about A.I. models’ overall capability growth, and yet that is precisely what its meager sample means it cannot do.

Jitse Goutbeek's avatar

Thank you for reciting this, I should have responded to this right away. I don't think the argument this paragraph makes is valid. If the distribution was completely spurious as you seem to assume based on this paragraph you would also expect it to not be predictive and for the results of AI-systems to be more or less random. Yet releases after the original result have been published have shown more or less the same pattern as the result was predicting, making a good predictive model based on a bunch of random numbers would be nearly impossible. And none of your criticisms suggest to me that these task time completions are random, just that they translate differently to real-world impact. If samples are too small to be meaningful and there is a lot of noise, you would expect the variance to be enormous and the results to not be predictive at all.

Not everything has to scale down by the same number for this to be meaningful. Maybe you cannot project it linearly on what you think it ought to be measuring, but that doesn't mean it isn't measuring something important that is correlated. The AI systems are clearly getting better at 'something' at an exponential rate.

You are saying this: "It is measuring whether A.I. can occasionally complete 97 highly contrived software engineering tasks whose “lengths” are spuriously determined. Nor does “extrapolating this trend” predict anything that can be understood in terms of what “humans,” as such, can do."

I don't think we can say their lengths are spuriously determined for the reasons above (or that the sample being small negates the result given predictive value), if they were there would not be a line to be drawn and AI completion rates would not be so correlated with them. In the worst case, they are determined by badly incentivized humans new to the task (btw the AI systems are also new to the tasks in this setting, and in real production they might be working on their own codebases as well). However, it is clear that even with the bad incentives their task time completion is quite correlated with what LLMs are and are not able to do, which suggests they still completed the easier tasks quicker than the hard ones and did try to actually complete the tasks. In the limit you would expect badly incentivized software engineers given enough time to still be able to complete tasks - after all real software engineers also don't always have perfect incentives and have once been new hires.

I also think we should expect tasks to be less contrived over time, it makes sense that at any moment in time if you try to measure AI-progress you are trying to find samples that allow you to differentiate the systems we have today - easy enough to complete sometimes hard enough to not be saturated. We have seen many times that as benchmarks get saturated new benchmarks emerge. If you look at the 50% messiest task thing you see increase but a lot of them are too close to 0 to say anything meaningful about the speed of this increase, this will change over time.

Nathan Witkin's avatar

I don’t think you’re understanding the argument in the paragraph I sent.

 >“If the distribution was completely spurious as you seem to assume based on this paragraph you would also expect it to not be predictive and for the results of AI-systems to be more or less random.”

This does not follow at all. My point is that the distribution of baseline times is not representative of the target population of ALL software engineers. It could still very well be predictive of how future models perform according to the times from METR’s small sample of baseliners. The key point is that you can’t generalize the slope of the resulting curve fit to those baselines, because they do not have sufficient statistical power to tell us anything about the analogous baseline times for software engineers generally. If you redid the baselines with a much larger, randomly samples population, you could get a very different-looking curve. Perhaps it would still be exponential, but that’s pure guesswork at this point given the evidence available.

(As a side note, the curve wasn't even predictive of the performance of more recent models; GPT 5 / Opus 4.5 improved well in excess of the original trend; you can see this in the first graph from the article).

TLDR the curve may very well be predictive of how future models perform by comparison to METR’s small sample of baseliners, but since you can’t generalize their completion times to other software engineers, neither can you generalize the slope of METR’s doubling time graph (since its slope is highly dependent on those specific baseline times).

At best, you can say something like “well, LLMs are probably getting better at software engineering tasks, or else no matter how small and biased the sample of baseliners you probably wouldn’t get an exponential curve.” That’s fine, but:

a. you cannot infer that, if they redid the baselines with a much larger random sample, the curve would still be exponential; and

b. even if you could infer that, we would know nothing about the precise shape of the curve, and so it wouldn’t be of much use in forecasting.

Jitse Goutbeek's avatar

Thank you for taking the time to respond again and the clarification. I think you are right that I did not properly understood what you were trying to say, and it makes more sense to me now.

I'll try to summarize what I understood from what you said in my own words, with an example to see if I understood it correctly:

There seems to be an exponential trend here, and it seems plausible that this increase is correlated with usefulness of software engineering in the real world but we have no idea how that correlation looks like.

Imagine we would want to know how fast we should expect world GDP to increase, and therefore we try to measure GDP each year and see the kind of relation we are getting. Now we don't have the resources to measure GDP in each country so instead we do it in a couple of countries were we have good data from, which happen to be just the Nordic countries for example. We do find that GDP seems to be increasing exponentially in these countries, and this holds after our initial study we keep measuring year after year and they seem to be getting richer at an (exponentially increasing) rate. What does this mean for world GDP?

It could be the case that the Nordics are just weird, we even have good reason to belief so it's no accident that these are precisely the countries we were able to collect data from. Maybe this relationship might continue holding within those countries, but the rest of the world is plagued by various problems that make their GDP have a lower slope, no exponential increase at all or maybe not even an increase. It would be impossible to predict from just this sample when poverty in the world ceases to exist.

A similar thing might be going on here, we have no idea how well things generalize outside of distribution and how long they hold. Therefore meaningless research right? It is not predictive of anything.

I think you are right that the METR research doesn't prove if and when LLMs can automate all software engineering tasks, and it was very useful for me to see you make this argument. That being said I still think it is massively useful research, because the results do make more sense under some theories of what is going on than others and your theory of what is going on with LLMs than others - and therefore can be useful evidence in favor or against a theory even if it doesn't 'prove' anything.

If you belief the world is not getting richer at all, you would need to come up with a good explanation of why the Nordics are getting richer so quickly and if you belief there is some fundamental compounding effect that makes it so that if GDP increases it tends to do so exponentially because your theory of economic growth includes compounding effects, seeing this trend in Nordic countries makes your theory more likely (even if it doesn't prove it). I think getting the METR graph results is much more likely if LLMs are getting better at software engineering and if this is correlated with total compute used (which will likely continue to increase exponentially for a while, and can cleanly explain the results) than if there is some fundamental limitation that makes it unlikely that LLMs are getting more useful for software engineering at all. Off course there are many other possible theories, explanations and data sources we need to consider (METR also had research where they showed a slowdown by people using LLMs for software engineering for example)

At the end of the day this evidence is not proof of anything, but any one piece of evidence rarely is. I like to look at all the data we have as something your theory needs to be able to explain, and theories that do that better are somewhat more likely to be predictive. This makes experimental research like this, even with small samples, still highly useful. The fact that there is a clear trend makes it hard to explain it away by mere variance.

Nathan Witkin's avatar

Of course - sorry if I was a bit terse before.

Your GDP example is spot-on, and I think I may go back in and add an example similar to this one to make the point clearer (I'm considering adding yours exactly, in which case I'd be sure to credit you).

To use your example: yes, finding GDP increases according to a certain trend in the Nordics would tell you something about the dynamics of GDP as such, but it would be very difficult to know precisely what without more data. And you certainly could not generalize from the Nordics' growth rate to the rest of the world (although, as you point out perhaps you could be relatively confident GDP was not on, say, a decreasing trend in roughly similar parts of the world).

I like your argument that, at a minimum, METR's results make more sense under some theories than others. I think there's something to this, but at the same time I think you're being too generous given how little data METR has, and how many sources of noise are affecting it (sample size, sample bias, poor incentive scheme, unrealistic tasks, etc.). With so much noise and so little data, METR's result are actually possible across a very wide range of worlds. In my view, the only worlds they rule out are ones in which AI is not improving at software engineering at all, or doing so very very slowly (in which case it'd be pretty weird to see exponential improvement in any sample, no matter how small or noisy). But since I think we already knew we were not in such a world, I have a hard time learning much from their results. It doesn't help that they present their results as generalizable to all software engineers (and even to all humans in some of their communications). I think we're in agreement now as to how misleading this is.

Elias Schmied's avatar

Thanks both of you, this was a great exchange that helped my understanding.

Catmint's avatar

GPT 5 and Opus 4.5 were both major improvements over what came before. Subjectively it felt like that was when AI went from being a waste of time at software to having some scattered uses here and there. That the METR graph correctly showed them off-trend is a point in favor, not against.

MP's avatar

1- The trend seems to be validated by accelerated amounts of customer money being spent on tools like Cursor and Claude Code.

2- The trend seems to be validated by endorsements of "AI-aided coding" by a large amount of high-profile computer programmers, including Linus Torvalds. On Twitter, people are all saying it doesn't make sense to write code by hand.

3- The trend seems to be validated by things like the Epoch Capabilities Index, with the acceleration in the METR Chart equating to an acceleration in the ECI.

4- The trend seems to be validated by AI achieving gold in the International Collegiate Programming Contest World Finals.

Carlos's avatar

AI-aided coding is pretty different from vibecoding, having the AI do everything. Linus Torvalds does not appear to have endorsed vibecoding (he says it's fine so long as it's not used for anything that matters, which is my own impression as a software developer too). I use Claude Code too, but it's too unreliable to be used in a fire and forget basis.

> Twitter chatter

Probably stuff coming from Silicon Valley. I always disliked Silicon Valley culture, I feel it's very irrational.

> International Collegiate Programming Contest World Finals.

I suspect AI will do quite well at algorithmic stuff (figuring out an algorithm is technically a small task), which is a pretty different skill from building a complex, production ready codebase without it turning into a pile of spaghetti. Guessing, but I think around 95% of real world software development is not about algorithms at all, more about figuring out requirements and managing complexity.

Austin Morrissey's avatar

Evaluating rapidly advancing technology is hard. The field is nascent and without doubt can be improved on many fronts. If we cannot benchmark performance between models and across domains, our ability to assess emerging threats is limited.

One limitation that precedes improving our methods is the dearth of people interested in this aspect of research. I read your post, and by taking the time to write it, it seems like this interests you. Have you considered designing evals?

Carlos's avatar

Maybe you couldn't write a paper with this eval, but I'm liking "can beat Pokemon Red" as a good evaluation of general reasoning capability (currently, it's looking like Claude Opus 4.5 cannot https://www.twitch.tv/claudeplayspokemon). In an interview, I also saw that the difficulty in evaluating AI capability is that it's capability is jagged, it's superhuman in some senses, then very sub-human in other important ways. I kinda suspect it can get quite good at coding but then just fail completely at understanding how the physical world works.

Duane McMullen's avatar

While these criticisms are well founded, I am more charitable about the existing value and potential of this particular benchmark.

AI needs benchmarks to measure progress. Even a flawed benchmark is better than no benchmark at all. This particular benchmark is well defined and documented, which makes it open to substantive criticism. That's a very good thing!

Despite the flaws, the benchmark allows comparison across AIs and remains convincing as a general demonstration of the AIs becoming increasingly capable. You don't have to believe the specific numbers the benchmark throws out to find credible its portrayal of consistent and rapid improvement and the reletive difference between AIs.

The 50% success criteria for the headline chart is because it is the best at showing progress at the current frontier. As this article compellingly argues, the 50% mark is starting to lose value as a measure of progress. The benchmark allows shifting to a higher success rate as the new headline benchmark. The crawling speed benchmark needs to be updated once the baby starts to walk.

Similarly, task completion times can be better defined by actual experts working on those tasks, rather than trained non-experts. That would be a viable further refinement of the model, generating useful benchmarks in expert task specific areas and more plausible aggregate scores across many tasks.

While the benchmark is flawed, it is better than the alternatives out there and has all the scaffolding necessary to become a better benchmark as, and if, AI advances.

Penn Griffey's avatar

I don’t believe this is a fair critique of the work. While the points raised by other commentators are good, there are two methodological defenses that I don't think have been raised yet:

1. The "Low-biased Sample Size" is actually appropriate for their goal.

You argue that sampling human completion times from a small group of engineers (three in some cases) makes for a poor estimate of "humanity’s" average speed. This is true. However, the authors aren't trying to build a universal human performance metric; they are building a predictive signal for LLM success.

To track progress, the metric doesn't need to perfectly represent the global population of engineers; it just needs to be an internally consistent predictor of success.

The logistic regressions show that these durations (even if biased toward a specific subgroup) predict model success quite well. If the goal is to track the trend in AI improvement relative to a fixed baseline, this "biased" baseline is a perfectly sound benchmark.

2. The 50% Threshold is the most statistically powerful choice.

Critiquing the 50% threshold as "too low" misses the mathematical reality of the logistic curve.

The magnitude of the derivative of a logistic function is maximal at the 50% point. This is the unique inflection point where the model is most sensitive to changes in the predictor (task duration).

When comparing different logistic curves against one another, the 50% threshold provides the highest statistical power. As you move toward 80% or 90%, you move into the "tails" of the distribution where the slope flattens.

Your observation that results "look worse" at 80% is actually expected: at higher thresholds, the signal is being "squashed" by the asymptotes of the curve, leading to less stable estimates. Using 50% isn't an arbitrary "low bar" it's a principled choice that captures the most robust part of the data.

Overall, while we can certainly debate the interpretation of what a 7-month doubling time means for the future of work, the methods used to identify that trend are statistically sound.

Nathan Witkin's avatar

I'm not going to respond to blatant AI slop. Try again in your own words.

akash's avatar

The point about incentives and possible inflation is quite interesting, but the “METR’s public comms choices are bad” seems unfair and wrong.

I was perpetually online when the RE-bench and HCAST papers were published, and my impression was that AI boosters would make exaggerated claims based on those charts but METR and researchers at METR were quite measured. One example: https://x.com/BethMayBarnes/status/1902759691727540329

The RE-bench paper is quite honest about its methodological limitations. HCAST references these limitations which you have cited in this article. I feel that is very strong evidence that METR is being transparent about how to carefully interpret these charts!

Nathan Witkin's avatar

I mean my criticisms re comms are of the charts themselves, as well as descriptions of them given by METR on their X and their official blog. Don't see how they're absolved by being more tentative in their informal comments (esp. given how few people see them). If anything, that makes it worse; why are you posting chart crimes you KNOW are chart crimes?

My sense is that they wanted to have their cake and eat it too, i.e. post heavily inflated / misleading results on their official accounts for the virality / discourse-shaping potential, while trying to save face on their private accounts. That strikes me as pretty objectionable.