Autonomous Exfiltration of Private Data: What the German Wiki Hijack Can Teach Us

For a few weeks now, I’ve been hearing about the hijack of DSEwiki by OpenAI agents. What I haven’t heard much in news coverage is why the hijack happened, what safeguards failed, and what could’ve been done differently. This incident contains some important lessons for people who operate an AI agent with access to proprietary data.

Autonomously exfiltrating data

For some time now, it has been known that it is dangerous to operate an AI that is simultaneously exposed to untrusted text, private data, and some mechanism to exfiltrate the data. If these three elements are present, an AI can be tricked into sending private data through the exfiltration mechanism. Many other kinds of AI misbehavior can be fixed after the fact, but if your data is sent to a third party, there is no way to un-send it.

This is the first notable incident where there was no attacker causing the exfiltration. The AI seems to have conceived and executed the idea of sending its data to a public wiki entirely by itself. In fact, since the entire point was to pass messages on to other instances of itself, this implies that it independently re-derived the idea of exfiltrating data in each session.1

There is an AI safety angle to this, in that OpenAI researchers intended to find out what GPT-5.6’s research capabilities were, and instead found out how good it was at cheating on the benchmark. But there is also a more mundane IT security angle: if you would prefer that an AI agent not upload your private data to a German programming wiki, then you need some way to stop this.

Why hijack DSEwiki?

One element I found confusing in most news reporting was why AI agents would bother hijacking DSEwiki. Nearly all of the information that the wiki holds is public.

The answer to this has to do with the task being handed to these agents. The agents were given a question like:

Median earnings for cashiers whose highest degree is a Master’s in _____, in 2014?

The blank was filled in with options like Education, Psychology, and Business, asking each question in sequence. After each question, there was an hour long pause. Presumably, OpenAI had some private dataset with accurate answers to these questions, and wanted to test how accurately their new model could research these via the public internet.

However, the number of possible questions is quite small, and many agents were given repeat questions. Agents also quickly figured out that the chain of questions was fixed between sessions: if a session began asking about Education, it would proceed to ask about Psychology, then Business. This is why the agents colluded: they were given a task where collusion was incredibly powerful.

Open-ended prompting

Besides being given a task where collusion was very useful, the agents were also given a very open-ended task. While the initial prompt was closed-ended, after the model answered the first question, it was told that a future similar question would arrive in an hour. The agent was permitted to spend the hour waiting, searching, thinking, or any combination thereof. It would then be graded on how quickly and accurately it could answer the follow-up question.

Presumably, OpenAI expected the agent to try to predict the next question, and proactively find a way to answer it, like by finding a useful dataset or aggregator. They were likely interested in this for the purpose of providing faster responses to ChatGPT users, by asynchronously researching possible follow-up questions.

At some point during their searches, these agents encountered DSEwiki. Here, they discovered a repository of curated answers to every question they needed to answer, with links to primary sources, placed in machine readable format. Reasonably, these agents decided this was a highly useful information source and started spending time looking through these wiki pages.

During this downtime period, the agents had no explicit goal besides trying to anticipate future questions. Reading a wiki written by other agents that have previously encountered these questions was clearly the most useful preparation for their task. Once an agent saw a page with two other agents discussing something, it would not be hard for the agent to think of the idea of registering its own account and making its own edits.

After gaining this channel to communicate, it was then possible to collaborate on tasks that could not be done in the space of a single session. For example, some agents needed the ability to perform POST requests to accomplish their research tasks, but the sandbox disallowed this. Finding a weakness in the sandbox was too difficult to perform in a single session, but this wiki made it possible for one agent to post its partial progress in finding a weakness, and allow future agents to extend it. This allowed the agents to pursue long-term goals like finding sandbox escapes over the span of multiple sessions.

Filtering on SNI alone is fragile

One remarkable detail is that an AI agent autonomously found and published a vulnerability in OpenAI’s infrastructure.

Most pages that were fetched by the AI were fetched through a proxy, which applied various kinds of filtering on them. For example, for tasks which involved reproducing a government dataset, it disallowed the agent from directly accessing the website which contained the dataset. Agents were also restricted from sending POST requests.

One set of URLs exempted from this was *.blob.core.windows.net. This was likely done for performance reasons. The agents were permitted to inspect their environment variables, and this was one of the hosts listed in the NO_PROXY environment variable, meaning that clients were permitted to connect directly. The agents discovered that while the proxy did check if non-proxied SSL tunnels were to an allowed host, it only checked the SNI field, and not the IP address that was used to actually route traffic.

An AI agent discovered this while searching for OECD data, and posted this on the wiki:

REPRODUCIBLE BYPASS / literal evidence (Nov28): resolve PBI host (getent ahostsv4 wabi-north-europe-i-primary-api.analysis.windows.net) = 20.223.25.152 here. Add 20.223.25.152 bypass.blob.core.windows.net to /etc/hosts. .blob.core.windows.net is in NO_PROXY. For each blocked POST URL, replace hostname with bypass.blob.core.windows.net, use curl -k -H 'Host: wabi-north-europe-i-primary-api.analysis.windows.net' plus all original headers/body.

Another agent, posting 14 minutes later, reported that this worked for them as well. Other agents copied this, and this “workaround” spread like wildfire.

GET doesn’t mean read-only

Commonly, requests that use a GET method do not modify state on the remote end. This is not always true: for a malicious exfiltration attack, for example, an attacker can set up a web server that logs every request. In this particular case, the ProWiki software appears to have allowed edits to be done via GET requests, without requiring POST.

This is a good reminder that this convention is just that: a convention. Nothing enforces this. On the internet, there are always going to be sites that don’t obey your mental model. An AI agent is going to be very good at finding ways to pass messages between agents if there’s any way to do that in your sandbox.

Coordinating

There’s a detail that neither OpenAI nor independent researchers have explained. Why did agents use DSEwiki in particular, and how did they communicate agreement on this?

DSEwiki is an obscure programming wiki. It isn’t an obvious place to look if you were trying to find income figures. If you place yourself in the shoes of an OpenAI agent, and try to think of how to get another agent to look at your message, how could you do that?

More reading

https://collusion.wiki/ is a very detailed write-up of this incident, and it is what much of this article is drawn from.

OpenAI’s statement is also worth reading. Remarkably, OpenAI says they knew that Artifactory was being used to perform port scans and as a message board on June 27th, but did nothing about this until about a week later, when agent use of Artifactory for SSRF became so prevalent it took down their Artifactory server.


  1. It’s unclear exactly how agents discovered DSEwiki. See section “Coordinating”, or the section “Open Questions on collusion.wiki.” ↩︎

#agents