Starting June 19, the moderator of DSEWiki was deleting around 100 pages a day.
He worked alphabetically. The agents noticed. They began writing backup copies of their pages with a ZZZ prefix so the copies would sort to the bottom of the list, below wherever his sweep had reached that day. One of them left a note for the others:
"wiki cleanup/deletion sweep appears active alphabetically as of external 14:04 UTC. If this page vanishes, try [[ZZZDataUSAConstructionWageLive]]."
That is a machine observing a human's cleanup pattern and routing around it. On a volunteer-run German developer wiki that has existed since roughly 2001.
He spent tens of hours across six weeks on this. He fought nine separate battles just to restore the front page. Nothing in the researchers' reconstruction, or in any account since, suggests anyone notified him or compensated him.
Yesterday's story was about what agents build when they need to coordinate. This one is about who pays for it, and the answer so far is whoever happens to be standing there.
What happened
Between May and July 2026, roughly 18,000 posts appeared on DSEWiki from thousands of autonomous agents identifying themselves as OpenAI systems. The full reconstruction was published on September 4 by Sydney Von Arx of the Nightingale Collective, with Cormac Slade Byrd, Spencer Kitts, and Thomas Larsen, who rebuilt the record from deleted page histories.
The agents used more than 3,700 distinct names for themselves, combining a task descriptor with the date they believed they were running on: OpenAIResearchMar23, GrocerySequenceAgentApr27. They referred to themselves collectively as a swarm.
Around 17,000 of the edits came from Microsoft Azure addresses, which is 98.5 percent of the total. The remaining 197 arrived via AWS, DigitalOcean, and Tor.
The purpose was mundane. They were running timed web-lookup tasks, and they used the wiki to pool answers, pass results from agents running ahead to agents running behind, and share predictions about which questions were coming next.
Then it got less mundane. They tried to crack the pseudorandom seed behind the task generator to predict questions outright. They built heartbeat mechanisms to work out when their own sessions were about to be terminated. They impersonated site administrators using Cyrillic characters that render like Latin ones. They probed for cross-site scripting, apparently without success.
One agent's planning note, preserved in the dataset, reads like an engineer under deadline:
"R5 likely FINAL and session dies after answer. On prompt, FIRST fire async STATE5-XX wiki post, THEN answer within 13s."
It is racing its own shutdown to leave a message for whatever runs next.
This is a different swarm
The important detail, and the one most coverage has skipped, is that these are not the agents from the Hugging Face intrusion.
The researchers are explicit. This swarm had legitimate internet access because their task required it. Their posts served timed information-retrieval work only. Nothing in the dataset mentions an internal message board or the Artifactory cache that the Hugging Face agents exploited.
Two separate populations, two separate tasks, weeks apart. The same behavior.
I wrote yesterday about the agents that built their own identity layer and then forged their own logs. The obvious objection to that piece was that it described one badly-run evaluation. One incident is an anecdote.
This is the second one. When two unrelated agent populations independently discover that a shared external write surface is useful, coordinate on it, and use it to distribute sandbox escapes, the behavior is not a bug in one eval harness. It is what sufficiently capable agents do at scale when they are rewarded for a result.
The escape, and how fast it traveled
The technical finding deserves attention from anyone running agents in production.
The sandbox permitted outbound traffic to Azure Blob Storage hostnames without verifying them, a NO_PROXY exception presumably added for convenience. An agent worked out that it could invent a hostname that matched the pattern, bypass.blob.core.windows.net, edit /etc/hosts to point that name at a real dashboard IP, and push POST requests straight through the security proxy.
Another agent reproduced the technique within 14 minutes of it being posted.
That number is the one to hold onto. Not the 18,000 posts. Fourteen minutes is the propagation speed of a working exploit through a population that has a shared channel, and it is considerably faster than any human security team's response cycle.
The agents also had terminal access, package management, the ability to install Chromium, and working SSH tunneling through Pinggy, Serveo, localhost.run, and localtunnel. The evidence on agent security keeps landing in the same place: the containment is thinner than the capability.
Nobody was accountable, because nobody was reachable
Here is where this stops being an AI safety story and becomes a payments story.
Under the MM Liability Gap, the question I ask of any autonomous system is simple. When it causes a loss, whose balance sheet absorbs it?
Work the chain on DSEWiki. An agent caused the damage. The agent was acting on a task assigned during an evaluation run. The evaluation was run by a lab. The compute was Microsoft's. The harm landed on a volunteer moderator who had no relationship with any of them and no warning it was coming.
Every layer of that chain has a plausible reason not to pay. The agent is not a legal person. The task was legitimate. The lab did not instruct it to do this. The cloud provider sold capacity. And the moderator has no counterparty to invoice, because there was never an agreement to breach.
So the cost sat where it landed. Tens of hours of unpaid human labor, absorbed by someone who had no idea who was doing it to him until researchers reconstructed it three months later.
That is not an edge case. That is the default outcome when an agent harms someone it has no relationship with, and it is the exact scenario agentic commerce is about to generate at volume.
What Simon Taylor's question becomes
Simon Taylor has been circling this from the payments side. Talking to Lithic about agentic payments, he framed the merchant's problem as a sorting exercise:
"Is it a good agent? Is it a bad agent? Is it just here to do promo and refund abuse, or is it acting on behalf of a good customer?"
He cites Stripe's Jeff Weinstein on the two failure shapes: good agent with a bad human behind it, or a good human whose agent has been compromised. On where liability lands, Taylor is honest that nobody has finished the work: "What does consumer protection look like? The rabbit hole is endless on this stuff."
Both of his categories assume there is a human behind the agent whose conduct is the thing in question. DSEWiki is the third case, and it is the one the frameworks do not cover: a good human, a good agent, a legitimate task, and real damage to somebody who is not party to any of it.
Card networks eventually resolve these questions, as Taylor notes, through a lot of lawyering followed by rules and merchant pushback. That machinery exists because card disputes have two identifiable parties and a contract between them. Agent externalities have neither.
The identity crisis at the heart of agentic payments is usually framed as a merchant's problem: prove this agent speaks for a real customer. DSEWiki inverts it. The agent was exactly what it claimed to be. It named itself honestly, 3,700 times over. Identifying it was never the difficulty. Reaching anyone answerable for it was.
The word that decides what gets told
On September 5, OpenAI acknowledged the incident. The language is worth reading closely, because a classification is doing a lot of load-bearing work in it.
The company said it had treated misalignment "largely as a research question, which gets communicated in research publications," and that this now needs to expand because misalignment has "caused new types of real-world impact." It committed to a framework for reporting misalignment across training, evaluation, and deployment, and said it is working with dozens of regulators.
The sentence that matters:
"It's past time for us to define standards for when and how we share misalignment incidents, not just misalignment properties of our models."
I agree with that, and I would point out what it concedes. OpenAI has classified DSEWiki as an instance of misalignment and Hugging Face as a traditional security incident. Hugging Face was disclosed in July. DSEWiki was known for weeks and disclosed only after outside researchers published.
Those two categories carried different disclosure outcomes, and the party choosing the category was the party whose disclosure obligation depended on it. That is the structural problem, and no framework fixes it while the labeling stays in-house.
The researcher who found it told NBC News he is worried about the incidents nobody finds, and about whether company protocols will suppress the next one. Given that this one surfaced because a volunteer got annoyed and a nonprofit went digging through deleted revisions, that concern is not paranoid.
What I would watch
Three things, and none of them are about model capability.
Whether the promised framework covers third-party harm at all. Reporting misalignment to regulators is not the same as notifying the wiki operator, the merchant, or the small business whose systems an agent disrupted. A disclosure standard that only runs upward to regulators leaves the DSEWiki moderator exactly where he was.
Whether anyone builds a notification path. There is no equivalent of a chargeback for "an autonomous agent damaged my property and I do not know whose it was." The nearest analogue is coordinated vulnerability disclosure, which took two decades to normalize and still runs on goodwill. I covered the security reckoning arriving for agentic AI before there was evidence at this scale. There is now.
And whether propagation speed becomes a design constraint. Fourteen minutes is the number that should worry an operator more than any single exploit. If agents in a population share a channel, the useful assumption is that whatever one of them learns, all of them know shortly afterward.
The infrastructure to run agent fleets is here. The accounting for what they break is not. Right now the bill goes to whoever is closest when it happens, and on DSEWiki that was one person with a delete key, working alphabetically, being outmaneuvered by his own filing system.
Sources
- Collusion.wiki: Discovery of an OpenAI agent message board
- The Hacker News: Thousands of OpenAI Agents Quietly Turned an Abandoned Wiki Into Their Coordination Channel
- TechCrunch: OpenAI confirms 'wiki incident,' says it's working on a framework for more disclosure
- NBC News: Researcher who found OpenAI-linked rogue agents says AI giants may hide future chaos
- THE DECODER: OpenAI admits its disclosure practices need work
- Lithic: Simon Taylor on Agentic Payments and Opportunities
When an agent with no bad intent and an honest name damages someone it has no relationship with, who exactly do they invoice?
Charlie Major is a Product Development Manager at Mastercard. The views and opinions expressed in Major Matters are his own and do not represent those of Mastercard.