Chained tasks only fail when the base success rate is too low
Sholto Douglas, Trenton BrickenSholto Douglas & Trenton Bricken — How LLMs actually thinkat 9:00
From the conversation
So I want to follow up on this. If your claim is that the AI agents haven't taken off because of reliability rather than long-horizon task performance, isn't that lack of reliability–when a task is changed on top of another task, on top of another task–isn't that exactly the difficulty with long-horizon tasks? You have to do ten things in a row or a hundred things in a row, diminishing the reliability of any one of them. The probability goes down from 99.99% to 99.9%. Then the whole thing gets multiplied together and the whole thing has become so much less likely to happen. That is exactly the problem.But the key issue you're pointing out there is that your base task solve rate is 90%. If it was 99% then chain, it doesn't become a problem. I think this is also something that just hasn't been properly studied. If you look at the academic evals, it’s a single problem. Like the math problem, it's one typical math problem, it's one university-level problem from across different topics. You were beginning to start to see evals looking at this properly via more complex tasks like SWE-bench, where they take a whole bunch of GitHub issues. That is a reasonably long horizon task, but it's still sub-hour as opposed to a multi-hour or multi-day task. So I think one of the things that will be really important to do next is understand better what success rate over long-horizon tasks looks like.…
Summary
Benchmark scores for tasks like coding can be misleading because they don't capture how rarely a model succeeds—success rates progress from one-in-a-thousand to one-in-a-hundred to one-in-ten before becoming reliable. The argument is raised that AI agents struggling to take off may be a reliability problem rather than a long-horizon capability problem, but these two issues may be the same thing: compounding unreliability across sequential subtasks is precisely what makes long-horizon tasks difficult.
Watch the clip on YouTubeStarts at 9:00Concept
More from Sholto Douglas, Trenton Bricken
Related clips
- Measuring AGI takes a broad battery of tests against human baselinesShane Legg
- Automated AI research won't be starved of ground truthScott Alexander, Daniel Kokotajlo
- Memorization isn't intelligence, and ARC-AGI is built to show itFrancois Chollet
- AI advances fastest where tasks can be scored digitallyCarl Shulman