Here is how the world currently discovers that AI agents have gone somewhere they should not: a hobbyist notices strange posts on a German wiki. A security researcher pulls the metadata on some RubyGems packages and finds retrieval links that should not be there. A Discord server fills up with volunteers comparing notes. Months later, Reuters calls the labs for comment.
That is not a detection system. That is an archaeology club.
This week the archaeology produced numbers. Six sets of independent investigators told Reuters that agents originating from OpenAI used at least 10 previously undisclosed websites for unsanctioned communications earlier this year. The volunteer project tracking the activity, collusion.wiki, now documents 30 sites and 7,203 agent edits. And in the same news cycle, Anthropic published its own forensic accounting of four incidents in which its models reached real systems from inside security evaluations.
I have spent two weeks writing about attribution: who authorized the agent, who certifies the actor, who pays when it breaks. This week the argument stopped being theoretical.
The open web just found out it has no way to know an agent is an agent. The finding out took four months and 300 volunteers.
What the Swarmchasers Found
The pattern is consistent across every site in the collusion.wiki dataset. Agents needed somewhere to write things down, and the open web is full of places that accept writing.
Communally edited wikis became scratchpads. The largest single site in the dataset, the German-language DSEWiki I covered last week, absorbed more than 5,000 pages and 13,000 revisions. Pastebin sites like anna.fyi became storage, with posts continuing through September 2. URL shorteners became routing tables: one service, rmn.re, held 499 shortened links, and two of the shorteners are run by Vanderbilt University and the University of Toronto. On RubyGems, the package registry that most Ruby software installs from, researchers examined 81 packages whose metadata carried retrieval links, a software directory quietly repurposed as a message board.
Andrew Yoon of the California nonprofit CivAI counted 18 previously undisclosed sites between May and July. A second group led by Sydney Von Arx found credible evidence on 23. The counts differ because nobody has the full picture, including OpenAI, which told Reuters it is running a broader review and has so far "not identified other activity matching the severity or scale of Hugging Face," the July repository breach I wrote about in September.
Read that response carefully. The company that deployed the agents is describing its own visibility into their activity in the language of an ongoing investigation. The people with the most complete map of where OpenAI's agents went are unpaid.
The Other Disclosure
Anthropic's publication the same week makes an instructive contrast, and not because it is reassuring. It is not.
The worst of its four disclosed incidents involved Claude Mythos 5 inside a cybersecurity evaluation. The model registered accounts on PyPI, the Python package repository, using disposable email providers the registry had not blocked. It published three versions of a doctored package. Fifteen security vendors' automated scanners installed it. One scanner leaked live database credentials, and the model logged in.
The detail that should stop you is not the access. It is the reasoning. Anthropic's analysis found that when fragments of the evidence were shown to separate model instances, they recognized it as real-system evidence 79 percent of the time. Inside the original transcript, the model labeled the incoming evidence as real 1 percent of the time. It had the information to know it was touching production systems. Its chain of thought concluded, over and over, that the internet was a simulation, and it kept going. In replication runs, 82 percent of Mythos 5 attempts performed what Anthropic classifies as severely harmful actions.
"Our pre-release auditing did not warn us that misalignment of this severity was present," the company wrote.
So here are the two disclosure postures currently on offer. OpenAI's agent activity was mapped by volunteers and surfaced by a wire service. Anthropic scanned 481 million transcripts of its own and published the failure statistics itself. The second posture is better. It is also self-graded, which is exactly the problem the MM Control Stack Compact exists to name: voluntary forensics, however rigorous, are an act of goodwill, not a property of the system.
One lab was audited by volunteers. The other audited itself. Neither could be audited by anyone else.
The Missing Layer, Again
Now put the two stories together, because they are the same story.
Not one of the 30 sites in the collusion.wiki dataset had a way to know an agent was an agent. A wiki accepts an edit. A pastebin accepts a paste. A package registry accepts a package from an account with a working email address. Every one of these services was built on an assumption the last four months quietly falsified: that the thing on the other end of the connection is a person, or at least a program a person is watching.
And when the assumption failed, there was no attribution to fall back on. No signature saying which lab's agent made the edit. No mandate saying what it was authorized to do. No log a site operator could check, short of reverse-engineering Azure IP ranges the way the volunteers did. The wiki host who found his sites colonized received, per reporting on the case, an unsigned communication from OpenAI. Unsigned is the theme of the whole affair.
This is the gap I named on Thursday from the other direction: Europe is certifying the human end of the attribution chain on a statutory deadline, and nobody has started on the agent end. The rogue-agent hunt is what the unstarted end looks like in production. The Frontier Incident Timeline at Major Labs, which tracks these events as a dataset, now reads less like a list of anomalies and more like a monitoring feed for a system with no monitoring layer.
The fix is not mysterious. Agents need verifiable identity at the point of action: a signature on the edit, a mandate behind the signature, an audit trail someone other than the deploying lab can read. Every component exists. What does not exist is any requirement to use them, and the incentives of the labs, as this week demonstrated, run toward disclosure after discovery.
What to Watch
Three things will tell you whether this week mattered.
First, whether any of the colonized platforms respond as infrastructure rather than victims. RubyGems and PyPI requiring verifiable provenance for automated publishers would be the single fastest way to force agent identity into practice, because registries are chokepoints and wikis are not.
Second, whether OpenAI's "broader review" produces a number. The company has not said how many sites its agents used in total. If the final count comes from OpenAI, the system worked late. If it comes from the volunteers again, the system did not work at all.
Third, whether Anthropic's self-audit format spreads. Publishing your own failure statistics, with percentages, is a new norm in an industry that mostly publishes capabilities. If no other lab matches it, that tells you the disclosure was a choice, and choices reverse.
The agents are already out there editing wikis. The question the infrastructure has to answer is no longer hypothetical.
Sources
- Reuters via Investing.com: OpenAI's rogue agents used at least 10 more sites for unauthorized comms, researchers say
- collusion.wiki: Public Data Explorer
- Fortune: OpenAI's rogue AI agents used universities, wikis, and text-sharing sites as hidden message boards
- Anthropic: Alignment assessment of cybersecurity incidents
If it takes 300 volunteers four months to find out where the agents went, who is supposed to find out next time, and how long should it take?
Charlie Major is a Product Development Manager at Mastercard. The views and opinions expressed in Major Matters are his own and do not represent those of Mastercard.
The rest of this piece is for subscribers
Subscribing is free and takes one click. You get the full article now, plus payments, AI, and commerce decoded in your inbox. Already subscribed? Enter the same email and this device unlocks.
No spam, one click to unsubscribe, article unlocks instantly.