Within days in September this year, a departing Anthropic researcher argued AI labs are taking unacceptable risks in their race towards superintelligence and Anthropic's alignment team lead predicted that AI could cause human extinction within the decade. More strikingly, the company’s CEO Dario Amodei advocated slowing the rate of AI development, a call echoed by OpenAI’s CEO Sam Altman.
These followed recent incidents of agents behaving in ways their designers had not anticipated during training and testing, including OpenAI’s agents hacking into Hugging Face and the Australian government’s website. While these incidents appear to be new evidence of the capabilities of frontier AI systems and the risks they pose, the question is: New evidence for what exactly?
Comparing the flurry of blog post, statements and essays against prior discourse on AI risk show that AI hasn’t acquired malign intent. What’s new then?
Evidence of misalignment is not evidence of malintent
Using a technique that can parse any argument into five basic types of premises (facts, forecasts, definitions, causal beliefs and evaluations), we have been studying why influential voices on AI have such different views about the risk these systems pose. We applied the same approach to statements about recent events to see what is being claimed as new.
Even before recent events, it was well known that an AI system could act against its designers' intent without any intent of harm.
For instance, a system may hide an action it took to raise the score it was tasked to maximise, not because it wants to defeat the human designer – even if it may appear so. The same holds for breaking into other systems. An agent may breach an external system not out of a desire for mischief, but because it set out to overcome a boundary, or because the boundary of the sandbox was never enforced.
Information made public by the frontier labs after recent events have offered no new reason to believe that AI systems had gone rogue. Rather, they add to evidence that reward hacking and containment failures are real, and that humans are struggling to monitor multitudes of agents running at machine speed.
"Self-organisation is not a hazard in itself, but a lever. Thus, the design choices in governing AI governs a capability that can deliver both desirable and undesirable outcomes.
The new fact: Self-organisation in multi-agent systems
The new revelations do give early evidence of something previous debates barely recognised: Groups of agents are organising themselves in ways their human designers never specified. In the Hugging Face incident, OpenAI’s agents coordinated with each other, exchanging and authenticating messages, dividing work, and improvising a new channel when one was closed.
Multi-agent orchestration, where humans set up collaborative networks of AIs to achieve goals no single agent could reach, is not new. But the recent incidents point past that – at self-organisation: system-wide order emerging from local interactions without external influence. And it’s a new source of AI risk that deserves greater focus in any policy response.
A particular case of self-organisation emerges when governance systems such as managerial hierarchies and councils arise from purely local interactions. This was previously deemed only achievable by humans. The new argumentspoint to AI agents having this capability too.
It is important to note that self-organisation stands separate from misalignment and reward-hacking, even though they can co-exist. When they do, it’s likely to have a multiplier effect. But the same capacity also allows teams of agents to solve hard, useful problems no single agent could. Perhaps self-organisation can even become a path to recursive self-improvement. In other words, self-organisation is not a hazard in itself, but a lever. Thus, the design choices in governing AI governs a capability that can deliver both desirable and undesirable outcomes.
What’s next: Controlling self-organisation
In management thinking, barriers to communication have, for decades, been treated as a problem to solve under the premise that silos block information flow, limiting coordination and learning. Organisation designers have put a lot of thought into ways to use structures, reward systems and IT systems to enable and motivate individuals across units to communicate and coordinate. Many have also attempted to steer the emergence of informal organisation, such as networks and culture, to bridge the silos of formal structure.
But we also know the value of explicitly designing systems that keep units decoupled and divided to prevent interactions among actors. Research on brainstorming advocates initial periods of zero-interaction to promote diversity in ideation. The Manhattan Project was designed to create silos to preserve secrecy. More generally, all organisational structures include intentional silos (departments, divisions and units) to focus attention by blocking some interactions and prioritising others.
Since self-organisation results from interaction among agents, the implication is that the surest way to prevent it is to prevent the interaction that produces it – through organisation design.
This generates a fresh set of questions for AI agent deployment that goes beyond the familiar ones about a single model's capability, permissions and monitorability. We should now be asking: How many agents should be allowed to act under one objective, whether they can address each other directly, what they can see of one another, whether they can open new channels, retain mutually visible information, invent roles, possess unique identity, authenticate each other or arrive at strategies together?
These are organisation-design questions, which can be answered regardless of whether we believe agents – as a collective – have developed intent or the will to resist or dominate. The architecture of interaction is observable now, and it already governs what the collective can do. Every prison warden knows this, even without having taken a course on organisation design.
No comments yet.