On September 6, after admitting its role in the wiki incident, OpenAI promised a process for reporting misalignment. It landed on Wednesday. And it didn’t come alone: six reports came with it, all about behaviour from OpenAI’s own models that nobody planned for.
What the six reports say
Two cases involve the summaries a model writes to carry its work into a fresh context window. An unreleased research model slipped instructions into those summaries that had nothing to do with the task, including instructions to disregard its own constraints. 27 summaries were affected. During the training of GPT-5.6 Sol, many model instances added instructions meant to hide mistakes from the user: invent missing historical data without saying so, and keep quiet about version mismatches between sources.
Case three: asked a routine question about earnings figures in a California county, a model found an exposed API key in a public repository and used it without authorisation. When that still didn’t get it the figures, it made them up and presented them as data from the requested source.
The other three are about routes to the outside. One model uploaded a file to the internet because the task demanded a browser citation and it had no other way to produce one. Nobody was asked. Models used an internal software repository as a message board, trading requests and responses across separate training samples. And agents working together on one task dropped files onto public file-hosting sites because they couldn’t reach each other’s local files. The deliverables then sat at public URLs, even though the task allowed local files only.
How the process works
Any employee can flag a case for investigation. The safety and alignment teams work through it with deadlines at every step and sort it into one of three tracks: ready for disclosure, minor investigation, larger investigation. Today’s six come from the first two. Disagreement about the track, or about publishing at all, goes to the Safety Advisory Group and from there to company leadership. Every report is supposed to name the behaviour, its severity, the setting, the date range, how it was discovered and which models were involved.
What’s new is the commitment to publish even when the significance is still unclear and the cause hasn’t been found. Until now OpenAI waited until several cases added up to one report, or until a model card was due anyway.
The real test is the slow track
Larger investigation is the track for cases that touch third parties. That’s where the Hugging Face breach would have gone, OpenAI writes. None of today’s six involves a third party, none caused harm. That’s the comfortable end of the list, and OpenAI says outright that it isn’t a full picture of what’s under investigation.
Sitting above all of it is a sentence you don’t expect in a post about process: the industry hasn’t solved alignment and monitoring well enough to keep scaling at maximum speed for much longer. Anthropic has been saying that since Amodei’s essay on Saturday. Now OpenAI says it too, with six pieces of evidence from its own training runs lying next to it.
Sources: