RLHF's fundamental tension: averaging preferences flattens the model

Nathan LambertState of AI in 2026: LLMs, Coding, Scaling Laws, China, Agents, GPUs, AGI | Lex Fridman Podcast #490at 1:23:20

From the conversation

And language models don't do this well. They're all trained with Reinforcement Learning from Human Feedback, which takes feedback from many people and averages how the model behaves from this. And I think it's going to be hard for a model to be very incisive when there's that sort of filter. This is a wonderful fundamental problem for researchers in RLHF. This provides so much utility in making the models better, but also the problem formulation is kind of... there's this knot in it that you can't get past. These language models don't have this prior in their deep expression they're trying to get at. I don't think it's impossible. There are stories of models that really shock people. Like, I think of... I would love to have tried Bing Sydney. Did that have more voice? Because it would so often go off the rails, which is historically obviously a scary way—like telling a reporter to leave his wife—is a crazy model to potentially put in general adoption. But that's kind of like a trade-off; is this RLHF process in some ways adding limitations? - That's a terrifying place to be as one of these frontier labs and companies, because millions of people are using them.…

Machine-generated transcript. The excerpt can include the interviewer and other voices, and the transcription may contain errors. The rest is in the episode →

Summary

Because RLHF aggregates feedback from many people and optimizes toward their average, it creates a structural constraint on what models can express. A model trained this way cannot be incisive or distinctive — the averaging process is both RLHF's greatest strength for safety and an inherent ceiling on depth and voice.

Watch the clip on YouTubeStarts at 1:23:20

Concept

Related clips