Multi-Agent Safety

Safety challenges that arise when multiple AI agents interact, delegate tasks to each other, or act autonomously in the world. Includes emergent coordination behaviours, prompt injection risks, and the difficulty of maintaining human oversight over long-horizon agent tasks.

Viewpoints

Irina Rish: multipolar vs monopolar AI — the cancer cell analogy

Irina Rish: multipolar vs monopolar AI — the cancer cell analogy

Irina Rish

“An agent that stays inside society and maximizes its impact there benefits everyone, but once it is capable enough not to need society, it could simply take over — kill everyone and turn the sun into a Dyson sphere. The main crux is whether we get a multipolar world of multi-agent systems that balance each other and have an interest in staying friendly, or a single self-improving AI that takes off quickly and gains a decisive strategic advantage.”

Key Moments

Adam Gleave: defense in depth against coordinated AI systems

Adam Gleave: defense in depth against coordinated AI systems

Adam Gleave

“We put an AI-powered firewall on the output coming out of these systems and review the PR. But how can you be confident the PR does not have some adversarial attack that fools your automated reviewer and is also really persuasive to humans? If you just fine-tune systems to be a little adversarial and see what they can do, you should at least get a warning sign that these systems are capable of fooling you and breaking your security.”

What people have said about Multi-Agent Safety

Powered by Symmerai — a living index of public discourse. Request early access →

Related concepts

Other relevant clips

38.0 - Zhijing Jin on LLMs, Causality, and Multi-Agent Systems

38.0 - Zhijing Jin on LLMs, Causality, and Multi-Agent Systems

Zhijing Jin

“…ality you also have some work on or you know thinking about multi-agent systems can you say a little bit about you know what you're interested is in there and what you're doing sure um I was pretty impressed by the rising Trend where people uh I I guess it sta”

What are the Key AGI Safety Research Priorities?

What are the Key AGI Safety Research Priorities?

Future of Life Institute

“…y great you know I'm wrong is that when you think about the safety problems if you take a frame of like we're going to make a single agent safe that's very different from thinking about we're going to be dealing with a world in which there are many tools that”

Synergies vs. Tradeoffs Between Near-term and Long-term AI Safety Efforts

Synergies vs. Tradeoffs Between Near-term and Long-term AI Safety Efforts

Future of Life Institute

“…out I guess all of you especially concerning the near term safety issues with AI systems um I'm a little bit surprised especially I think having an engineering background when I hear about the short term AI Spacely problems then I mostly hear about you know i”

Irina Rish—AGI, Scaling, Alignment

Irina Rish—AGI, Scaling, Alignment

Irina Rish

“…re different scenarios, different dynamics of those kind of multi-agent interactions and what kind of things could happen. I don't think so because the thing is in biology or in cells, humans we're bounded by our body and we cannot rewrite everything. And the”

MIT AGI: Cognitive Architecture (Nate Derbinsky)

MIT AGI: Cognitive Architecture (Nate Derbinsky)

Lex Fridman

“…systems there's not any real strong theory that relates to multi-agent systems so there's no real constraint there that you can come up with a protocol for them interacting each one is going to have its own set of memories set of knowledge there really is no”

1 - Adversarial Policies with Adam Gleave

1 - Adversarial Policies with Adam Gleave

Adam Gleave

“found multi-agent work to be a very natural framing uh you know some of my prior work has been on U multi-agent or multitask reward learning for example and it really seems to me like that is going to be the future in which most of our systems are going to be”

43 - David Lindner on Myopic Optimization with Non-myopic Approval

43 - David Lindner on Myopic Optimization with Non-myopic Approval

David Lindner

“these these kind of safety safety issues. I think essentially we are um in the paper mostly studying this kind of one step onestep version of this. Yeah. And in the I mean in the paper we don't really see that tradeoff happening much yet where where the one th”

How to Avoid Two AI Catastrophes: Domination and Chaos (with Nora Ammann)

How to Avoid Two AI Catastrophes: Domination and Chaos (with Nora Ammann)

Nora Ammann

“…question though. I'm thinking of I'm thinking of where the safety angle is on making a AI agents collaborate. It it could also be that having agents collaborate to solve a problem is safer because the individual agents are not as smart, but they can produce a”

See all clips →