Chained tasks only fail when the base success rate is too low

Sholto Douglas, Trenton BrickenSholto Douglas & Trenton Bricken — How LLMs actually thinkat 9:00

From the conversation

So I want to follow up on this. If your claim is that the AI agents haven't taken off because of reliability rather than long-horizon task performance, isn't that lack of reliability–when a task is changed on top of another task, on top of another task–isn't that exactly the difficulty with long-horizon tasks? You have to do ten things in a row or a hundred things in a row, diminishing the reliability of any one of them. The probability goes down from 99.99% to 99.9%. Then the whole thing gets multiplied together and the whole thing has become so much less likely to happen. That is exactly the problem.But the key issue you're pointing out there is that your base task solve rate is 90%. If it was 99% then chain, it doesn't become a problem. I think this is also something that just hasn't been properly studied. If you look at the academic evals, it’s a single problem. Like the math problem, it's one typical math problem, it's one university-level problem from across different topics. You were beginning to start to see evals looking at this properly via more complex tasks like SWE-bench, where they take a whole bunch of GitHub issues. That is a reasonably long horizon task, but it's still sub-hour as opposed to a multi-hour or multi-day task. So I think one of the things that will be really important to do next is understand better what success rate over long-horizon tasks looks like.…

Machine-generated transcript. The excerpt can include the interviewer and other voices, and the transcription may contain errors. The rest is in the episode →

Summary

Benchmark scores for tasks like coding can be misleading because they don't capture how rarely a model succeeds—success rates progress from one-in-a-thousand to one-in-a-hundred to one-in-ten before becoming reliable. The argument is raised that AI agents struggling to take off may be a reliability problem rather than a long-horizon capability problem, but these two issues may be the same thing: compounding unreliability across sequential subtasks is precisely what makes long-horizon tasks difficult.

Watch the clip on YouTubeStarts at 9:00

Concept

More from Sholto Douglas, Trenton Bricken

Related clips