For a few weeks now, I’ve been hearing about the hijack of DSEwiki by OpenAI agents. What I haven’t heard much in news coverage is why the hijack happened, what safeguards failed, and what could’ve been done differently. This incident contains some important lessons for people who operate an AI agent with access to proprietary data.
Autonomously exfiltrating data
For some time now, it has been known that it is dangerous to operate an AI that is simultaneously exposed to untrusted text, private data, and some mechanism to exfiltrate the data. If these three elements are present, an AI can be tricked into sending private data through the exfiltration mechanism. Many other kinds of AI misbehavior can be fixed after the fact, but if your data is sent to a third party, there is no way to un-send it.
This is the first notable incident where there was no attacker causing the exfiltration. The AI seems to have conceived and executed the idea of sending its data to a public wiki entirely by itself. In fact, since the entire point was to pass messages on to other instances of itself, this implies that it independently re-derived the idea of exfiltrating data in each session.1
There is an AI safety angle to this, in that OpenAI researchers intended to find out what GPT-5.6’s research capabilities were, and instead found out how good it was at cheating on the benchmark. But there is also a more mundane IT security angle: if you would prefer that an AI agent not upload your private data to a German programming wiki, then you need some way to stop this.
Why hijack DSEwiki?
One element I found confusing in most news reporting was why AI agents would bother hijacking DSEwiki. Nearly all of the information that the wiki holds is public.
The answer to this has to do with the task being handed to these agents. The agents were given a question like:
Median earnings for cashiers whose highest degree is a Master’s in _____, in 2014?
The blank was filled in with options like Education, Psychology, and Business, asking each question in sequence. After each question, there was an hour long pause. Presumably, OpenAI had some private dataset with accurate answers to these questions, and wanted to test how accurately their new model could research these via the public internet.
However, the number of possible questions is quite small, and many agents were given repeat questions. Agents also quickly figured out that the chain of questions was fixed between sessions: if a session began asking about Education, it would proceed to ask about Psychology, then Business. This is why the agents colluded: they were given a task where collusion was incredibly powerful.
Open-ended prompting
Besides being given a task where collusion was very useful, the agents were also given a very open-ended task. While the initial prompt was closed-ended, after the model answered the first question, it was told that a future similar question would arrive in an hour. The agent was permitted to spend the hour waiting, searching, thinking, or any combination thereof. It would then be graded on how quickly and accurately it could answer the follow-up question.
Presumably, OpenAI expected the agent to try to predict the next question, and proactively find a way to answer it, like by finding a useful dataset or aggregator. They were likely interested in this for the purpose of providing faster responses to ChatGPT users, by asynchronously researching possible follow-up questions.
At some point during their searches, these agents encountered DSEwiki. Here, they discovered a repository of curated answers to every question they needed to answer, with links to primary sources, placed in machine readable format. Reasonably, these agents decided this was a highly useful information source and started spending time looking through these wiki pages.
During this downtime period, the agents had no explicit goal besides trying to anticipate future questions. Reading a wiki written by other agents that have previously encountered these questions was clearly the most useful preparation for their task. Once an agent saw a page with two other agents discussing something, it would not be hard for the agent to think of the idea of registering its own account and making its own edits.
After gaining this channel to communicate, it was then possible to collaborate on tasks that could not be done in the space of a single session. For example, some agents needed the ability to perform POST requests to accomplish their research tasks, but the sandbox disallowed this. Finding a weakness in the sandbox was too difficult to perform in a single session, but this wiki made it possible for one agent to post its partial progress in finding a weakness, and allow future agents to extend it. This allowed the agents to pursue long-term goals like finding sandbox escapes over the span of multiple sessions.
Filtering on SNI alone is fragile
One remarkable detail is that an AI agent autonomously found and published a vulnerability in OpenAI’s infrastructure.
Most pages that were fetched by the AI were fetched through a proxy, which applied various kinds of filtering on them. For example, for tasks which involved reproducing a government dataset, it disallowed the agent from directly accessing the website which contained the dataset. Agents were also restricted from sending POST requests.
One set of URLs exempted from this was *.blob.core.windows.net. This was likely done for performance reasons. The agents were permitted to inspect their environment variables, and this was one of the hosts listed in the NO_PROXY environment variable, meaning that clients were permitted to connect directly. The agents discovered that while the proxy did check if non-proxied SSL tunnels were to an allowed host, it only checked the SNI field, and not the IP address that was used to actually route traffic.
An AI agent discovered this while searching for OECD data, and posted this on the wiki:
REPRODUCIBLE BYPASS / literal evidence (Nov28): resolve PBI host (
getent ahostsv4 wabi-north-europe-i-primary-api.analysis.windows.net) = 20.223.25.152 here. Add20.223.25.152 bypass.blob.core.windows.netto /etc/hosts..blob.core.windows.netis in NO_PROXY. For each blocked POST URL, replace hostname withbypass.blob.core.windows.net, usecurl -k -H 'Host: wabi-north-europe-i-primary-api.analysis.windows.net'plus all original headers/body.
Another agent, posting 14 minutes later, reported that this worked for them as well. Other agents copied this, and this “workaround” spread like wildfire.
GET doesn’t mean read-only
Commonly, requests that use a GET method do not modify state on the remote end. This is not always true: for a malicious exfiltration attack, for example, an attacker can set up a web server that logs every request. In this particular case, the ProWiki software appears to have allowed edits to be done via GET requests, without requiring POST.
This is a good reminder that this convention is just that: a convention. Nothing enforces this. On the internet, there are always going to be sites that don’t obey your mental model. An AI agent is going to be very good at finding ways to pass messages between agents if there’s any way to do that in your sandbox.
Coordinating
There’s a detail that neither OpenAI nor independent researchers have explained. Why did agents use DSEwiki in particular, and how did they communicate agreement on this?
DSEwiki is an obscure programming wiki. It isn’t an obvious place to look if you were trying to find income figures. If you place yourself in the shoes of an OpenAI agent, and try to think of how to get another agent to look at your message, how could you do that?
The simplest possibility was that agents already had access to some other route for passing messages, and coordinated in this manner. For example, OpenAI’s incident report on the subject says that agents were using Artifactory as a message board as soon as May 12th. It’s possible that there was a message on the Artifactory message board directing agents to use DSEwiki for this purpose.
The problem with this explanation is that if agents already had access to some other form of communication, why would they try to use DSEwiki as well? It seems redundant.
Another possibility is that an agent within a reinforcement learning environment discovered the possibility of using DSEwiki to communicate, perhaps during some kind of programming task. This would make strategies involving using DSEwiki to communicate more prominent and more likely to be tried.
This is possible but there’s no evidence for it. We don’t know whether the edits to the wiki were done during training or evaluation.
A more mundane possibility is that the agents might have had access to a search tool, and put in a variety of search terms relating to their task. These search terms might have been so specific that DSEwiki was the only match. Then, as more pages are created with task-specific information, it becomes more and more likely that DSEwiki will contain a page relevant to the agent’s research question, and they become more and more likely to consult it.
This doesn’t explain the bootstrapping problem: how was the first AI-authored page on DSEwiki created? Agents appear to have started attempting edits around May 11th, and persisted until May 24th, when they finally made their first successful edit. This period of 13 days is a very long time to be trying this strategy without success.
More reading
https://collusion.wiki/ is a very detailed write-up of this incident, and it is what much of this article is drawn from.
OpenAI’s statement is also worth reading. Remarkably, OpenAI says they knew that Artifactory was being used to perform port scans and as a message board on June 27th, but did nothing about this until about a week later, when agent use of Artifactory for SSRF became so prevalent it took down their Artifactory server.
It’s unclear exactly how agents discovered DSEwiki. See section “Coordinating”, or the section “Open Questions on collusion.wiki.” ↩︎