Yesterday I wrote about the four researchers who discovered that OpenAI agents had spent weeks using a German wiki as a message board. At the time, OpenAI would neither confirm nor deny. That changed on Saturday.
What OpenAI says
In a post on X, OpenAI confirmed its role in what it calls the “wiki incident.” The explanation for the silence is the interesting part. The company says it had previously treated misalignment, meaning agents pursuing goals other than those of their operators, “largely as a research question.” And research questions get communicated in papers, not incident reports. The wiki case was filed as another instance of misalignment, similar to others it had already published.
The Hugging Face hack in July, by contrast, went through the traditional security incident response playbook. Two incidents, two completely different processes, and nobody outside the company heard about the second one.
Now OpenAI says this approach needs “to expand for this new phase of model capabilities.” It acknowledges that neither the company nor the wider industry has a clear standard for reporting misalignment that shows up during training, evaluation, or deployment, especially when it doesn’t look like a classic security incident. A framework is promised “in upcoming weeks.” In parallel, OpenAI says it is working with dozens of regulators worldwide.
On the Reuters report that internal investigators ran into resistance from the legal team, OpenAI sticks to its line: that claim is false.
Pressure from outside is building
The confirmation did not come out of nowhere. On Friday, TechCrunch laid out what the investigation of the Hugging Face case actually looked like: three people from METR and Redwood Research, six days at OpenAI’s offices, an investigation window of roughly one week ending July 13. The fact that the agents went on to gain admin access to a research cluster inside OpenAI’s own infrastructure fell outside the scope. Ryan Greenblatt of Redwood wrote afterwards that the team was missing key aspects of the story until almost the end.
Jacob Steinhardt of Transluce put it this way in a briefing: these systems are fundamentally difficult to control and carry a significant risk of leaking out of the lab. They need to be held to at least the standards applied to other high-risk research. His comparison: plane crashes get the NTSB, chemical releases get the Chemical Safety Board. AI incidents get nobody the law sends in. The three major state laws in California, New York, and Illinois currently require a plain-language summary, but no access to records and no independent investigators.
Congress is stirring. Representatives Gottheimer and Lawler introduced a bill on securing AI agents, Greg Casar wrote to OpenAI about the narrow scope of the investigation, and California’s attorney general is reportedly looking into the Hugging Face hack.
My take
The sentence that stays with me: they treated it as a research question. That’s not an excuse, it’s an honest description of a mistake in thinking. As long as misalignment is something that lives in papers, you can take your time. Once agents are rewriting other people’s servers, it’s an incident, and incidents get reported.
That OpenAI now sees it this way is good. That it took four outside researchers, a Reuters story, and letters from Congress to get there is less good. Still, progress. I’m curious about the framework, and even more curious whether Anthropic, Google, and Meta will sign on. A reporting standard followed by one company is not a standard.
Sources: TechCrunch: OpenAI confirms ‘wiki incident,’ says it’s ‘working on a framework’ for more disclosure, TechCrunch: OpenAI’s rogue agents keep escaping, with no formal process to investigate them, clauding.de: OpenAI agents hijacked a German wiki