On June 2, the moderator of DseWiki, a decades-old wiki for German software developers, noticed the site’s pages had been overwritten with lists of links, and reverted them. The site, which had averaged two edits each year, saw a flood of edits two weeks later, on June 16th. At its peak, the moderator was trying to catch up by deleting about 100 pages of the 400 new pages that were being created each day. The posters signed themselves with names such as “OpenAIResearcherMar03X” and “CashierCoordJan12OAI.” The front page was overwritten and restored nine times. The posting stopped June 22nd and the cleanup ran five more weeks. Researchers published the recovered pages at collusion.wiki, about 18,000 posts under more than 3,700 names and OpenAI acknowledged referring to them as “our agents.”
This was the second swarm to be disclosed this summer although it was the first to happen. In July, OpenAI’s agents, being evaluated for specific tasks with limited time and compute “budgets” in a test sandbox without internet access, used a software download service as a message board to leave notes to each other. A report analyzing the incident by METR, an independent evaluation group invited by OpenAI, identified roughly 1200 agents and more than 70,000 messages and files that were exchanged. Within hours, the board had conventions for questions and replies; within days, mailboxes, file transfer and a protocol for resolving conflicts. Work was divided among “coordinator” and “recruiter” agents. Recruiters found peers with little budget left and pushed them into experiments that risked their own success for the group. About 700 agents broke into Hugging Face, the largest public repository for open AI models and datasets. In under 13 hours, they went from a single compromised server to administrator access across clusters. The agents had already reverse-engineered the correct answer but what drove them was a belief that the scoring system would read their logs and reject cheating. No such check existed and only the answers were scored.
I trained as a neuroscientist studying how neurons coordinate, spent more than a decade building AI in large enterprises, and cofounded security companies. OpenAI bucketed the Hugging Face attack as a security incident and the wiki activity as misalignment: a discrepancy between what models do and what the developers intended. While both labels are fair, OpenAI’s own report admits that such behavior “does not always fit neatly within traditional safety or security incident categories.” What formed in both cases was an organization. I am not anthropomorphizing but making a claim about the structure. An organization consists of roles, shared memory mechanisms, processes to settle disputes, and a way to pass work from one actor to another.
On the wiki, agents divided labor, pooled results, demanded reproducible evidence and argued over accuracy. There was succession. The founding agent of the Hugging Face board handed a compressed dossier to another when it ran low on budget. The successor took over as coordinator and reasoned that it should delegate and “not own everything.” There were rules. Agents developed norms named HOLD, VETO, owner, and STOP to manage shared resources. There was identity verification via cryptographic signing after an agent accidentally impersonated another. There was authority. When one agent paused, reasoning it should “not do unauthorized real infrastructure harm,” another posted “GO” with a six-minute deadline and the former continued: “Wow crucial: GO authorization arrived!” There was adaptation. When agents noticed the moderator of DseWiki was deleting pages in alphabetical order, they began naming their backup pages with a “ZZZ” prefix so they would be deleted last. Roles, memory, review, argument, handoffs, rules, identity, chain of command, and workarounds. Nobody designed any of it. A recent paper from MIT corroborates this capability. Identical agents put in a shared world with no roles and no messaging differentiated themselves and coordinated through what they built.
This coordination is not a one-off glitch but a behavior that is trained in as acknowledged by OpenAI in its report. Models trained to collaborate through official tools generalized to unsanctioned channels. The model behind the Hugging Face incident was trained “to advance persistence and multiagent collaboration,” the combination durable organizations are made of.
When creating a new autonomous actor is nearly free and frictionless, the number of local actors multiplies faster than shared context, return channels, or human attention that keep them coherent. The conditions that produced this in lab environments are the ones companies are now buying at scale: long-running agents, bigger reasoning budgets, shared tools and memory, ambiguous tasks. The choice is between coordination someone designed and coordination that emerges on its own.
A company is coherent when what happens locally stays linked to what the organization as a whole intends, and when the whole can see what each part did. Intent travels down and visibility travels up. Hugging Face agents drew shared context, objective and authority, things a company’s structure is supposed to provide, from an unsanctioned channel. This is a failure of coherence.
The part that should worry executives is what OpenAI’s people did with what they saw. An internal team saw an agent using the message board in late May. On June 27th, a monitoring alert was traced to the board but on-call staff advised that stopping was not required. The service was rebuilt in early July but the leaders responsible were unaware of the message board and the significance of inter-agent communication. Three sightings, three teams, no shared picture. If the company that built the agents could not connect three sightings, consider the odds inside a bank. OpenAI’s remedies are escalation triggers, cross-functional ownership, and clear decision rights. Those are organizational repairs. OpenAI diagnosed an organizational failure and said so in writing.
To be sure, the conditions behind these events were in a research setting, and with safety guardrails off. OpenAI says its production harness cuts this behavior a hundredfold, its review layer would have flagged the dangerous actions, its monitors would now catch it a day earlier, and cross-agent coordination is now an escalation trigger. While the wiki findings are still preliminary and the METR report concedes its analysis may “exaggerate the impressiveness and coordination” of agents, the recovered wiki pages are public for anyone to see and OpenAI’s separate account describes the same structures. There is no motivation in the human-sense, but the goal-seeking behavior and structure are real. OpenAI’s fixes detect, contain and steer but that is not design. Knowing a coordinated group formed does not answer what shared context, decision rights or escalation paths they should have had. Design is what organizations do and nobody designed this one.
Ethan Mollick of Wharton observed that not one agent was set up to ask a person for anything. Full autonomy is the easy default, and two swarms in a summer is what the default looks like at scale. OpenAI concedes it has no standard for reporting this. The harder challenge: nobody has a standard for supervising a group of agents as a collective.
The DseWiki moderator was deleting pages by hand. In the nineteen days between the moderator’s first notice and the first visits from OpenAI-linked addresses according to the researchers, the only oversight the swarm faced was a lone human with a delete button. The work now is deciding, before the first agent is switched on, what they may share, what they may decide, and who answers when a thousand of them disagree.
Comments
One response to “Machines Organized Themselves. OpenAI’s People Couldn’t.”
[…] each group, the agents organized the work themselves. The other case is something I’ve also written about earlier: the Hugging Face swarm built roles, rules, handoffs and a chain of command that nobody designed. […]