Discussion about this post

User's avatar
Keller Scholl's avatar

You're far too harsh:

1. On a fixed budget, I think a higher number of tasks, even with a lower number of participant-completions, was the right move. Yes, they've got very few, but 1 would be non-zero informative and 3 is in fact something. And I'd much rather 3 people doing 100 tasks than 100 people doing 3 tasks.

2. 50% reliability seems reasonable if your plan is "set one of my 5-8 simultaneously running agents off to do something, check back with it in a few hours". And that seems to be the workflow of many AI users I know. So long as the task is checkable...

3. One of the benefits of recruiting from people in your social network, rather than randoms you're paying, is that they're not going to deliberately slow their work, and you have some evidence of competence.

Elias Schmied's avatar

Useful, thanks (although a bit harsh for my liking).

"I expect some of you will want to say something like: “well, even if the specific times are off, the METR graph still shows model capabilities doubling at a frightening rate.” No. Do not try this. If the specific times are wrong, so are their distribution. That is to say: were we to collect baselines from a much larger, more representative sample, some tasks might fall into different completion-time buckets. For instance, tasks METR found took 2-3 hours might turn out, in this larger sample, to take 1-1.5 hrs. But, unless all completion times shift downward by the exact same proportion, this would compress the overall distribution of task completion times. That would in turn change the rate at which A.I. models moved up METR’s y-axis."

I'd have liked to see more elaboration on this point, since it's the most important one in the whole post. I'm not sure I have a good grasp of how strong this argument is.

Edit: ah, your exchange with Jitse in the comments does some of this.

24 more comments...

No posts

Ready for more?