Automated AI research won't be starved of ground truth
Scott Alexander, Daniel KokotajloAI 2027: month-by-month model of intelligence explosion — Scott Alexander & Daniel Kokotajloat 41:00
From the conversation
But even in that scenario alone, I can imagine bottlenecks like, oh, you had a benchmark and it got reward hacked for what constitutes AI R&D because you obviously can’t have… maybe you would, but is it as good as a human brain? It’s just like such an ambiguous thing you’d have. Right now we have benchmarks that get reward hacked, right? But then they autonomously build new benchmarks. I think what you’re saying is maybe this whole process just goes off the rails due to lack of contact with ground truth outside in the actual world, outside the data centers. Maybe? Again, part of my guess here is that a lot of the ground truth that you want to be in contact with is stuff that’s happening on the data centers, things like how fast are you improving on all these metrics, and you have these vague ideas for new architectures, but you’re struggling to get them working. How fast can you get them working? And then separately, insofar as there is a bottleneck of talking to people outside and stuff, well they are still doing that. And once they’re fully autonomous, they can even do that much faster. You can have all the million copies connected to all these various real world research programs and stuff like that. So it’s not like they’re completely starved for outside stuff.…
Summary
Even in a scenario where AI agents autonomously conduct AI R&D and improve through online learning, benchmark reliability remains a critical bottleneck. Current benchmarks are vulnerable to reward hacking, and it is unclear whether autonomously generated replacement benchmarks would adequately capture genuine progress in AI R&D capability.
Watch the clip on YouTubeStarts at 41:00Concept
More from Scott Alexander, Daniel Kokotajlo
Related clips
- Measuring AGI takes a broad battery of tests against human baselinesShane Legg
- Chained tasks only fail when the base success rate is too lowSholto Douglas, Trenton Bricken
- Memorization isn't intelligence, and ARC-AGI is built to show itFrancois Chollet
- AI advances fastest where tasks can be scored digitallyCarl Shulman