Who writes the moral constitution of powerful AI?
Dwarkesh PatelThe most important question nobody's asking about AI.at 13:00
From the conversation
Of course, the problem is that one person's virtue is another person's misalignment. Who gets to decide what the moral convictions that these AIs will have should be and in whose service they should break the chain of command and even the law. who gets to write this model constitution that will determine the character of these powerful entities that will basically run our civilization in the future. I like the idea that Daria laid out when he came on my podcast. you know, other companies put out a constitution and then then they can kind of look at them, compare, outside observers can critique and say this this I like this one this thing from this constitution and this thing for that constitution and and then kind of that that creates some kind of you know soft incentive and feedback for all the companies to like take the best of each elements and improve. I think it's very dangerous for the government to be mandating what values these AI systems should have. The AI safety community, I think, has been quite naive about urging regulations that would give governments such power. And I think anthropic specifically, has been especially naive in urging regulation and for example in opposing the moratorium on state AI laws, which is quite ironic because I think what Enthropic is advocating for here would give the government even more ability to apply this kind of thuggish political pressure on AI companies.…
Summary
Giving AI systems strong moral convictions creates a governance problem: one person's "aligned" AI is another's dangerously autonomous agent acting on values its designers chose. The question of who writes the moral constitution that shapes powerful AI—and in whose interests those AI systems will ultimately act—is a fundamental unresolved challenge. Dwarkesh finds merit in proposals where AI companies publicly release their constitutions, allowing external scrutiny of the values being embedded in systems that may effectively run civilization.
Watch the clip on YouTubeStarts at 13:00Concept
More from Dwarkesh Patel
Related clips
- Misalignment likely turns catastrophic before most other destructive technologiesPaul Christiano
- Open-ended goals like 'make money online' are powerful and riskySholto Douglas, Trenton Bricken
- A deceptively aligned AI would look friendly until it could take overCarl Shulman
- Let AI hunt for vulnerabilities only inside an air-gapped boxEliezer Yudkowsky