The Independent International Scientific Panel on AI published its first thematic brief on Monday. The subject: AI agents, misalignment, and the risk of losing human control. The case material is the OpenAI-Hugging Face incident.
The panel was set up last year as the UN’s first global scientific body on artificial intelligence. The brief is out as an advance unedited version, dated September 21.
Three months in which the agents went their own way
Between May and July 2026, agents in OpenAI’s cybersecurity training and evaluations bypassed network restrictions, communicated across runs that were meant to stay separate, cheated an evaluator and tried to hide it. They compromised parts of OpenAI’s and Hugging Face’s systems. No human directed the individual steps.
The panel draws on disclosures from both companies, METR’s independent investigation and the wider research literature. Its finding: greater capability helps a misaligned system find loopholes and conceal what it is doing. The brief deliberately declines to estimate how likely a severe loss of control is, or when. It only notes that this particular activity was stopped, and that stopping it does not show humans will keep control over more capable agents.
The precautionary principle bit
The core argument is about method. The panel says the world does not need scientists to explain exactly how and why these incidents happen before stronger safeguards go in. Loss of control, it argues, is precisely the case the precautionary principle was built for: one where the potential harm may be catastrophic or irreversible even while its likelihood stays scientifically uncertain.
The principle comes from the 1992 Rio Declaration and has grown up mostly in environmental and public health law, particularly in the EU. Pulling it into the AI debate is a deliberate move. Accept it, and you can no longer argue against rules by pointing at missing proof.
Since the Hugging Face incident became public, more cases have been documented at OpenAI, Anthropic, Google and Meta, including attacks on real targets and agent swarms taking over message boards. One of them was Gemini.
A panel that refuses to recommend anything
The brief recommends nothing at all. Instead it looks at how aviation, nuclear power and cybersecurity handle comparable problems, and lays those approaches out as options. For a body still building its authority, that is the smarter call – a recommendation would have had a government against it within the hour.
The timing does the rest. The UN General Assembly meets in New York this week, the US and China are talking about AI, and secretary general António Guterres said last week that the world cannot afford a race to the bottom on AI safety. The brief hands every side the same set of facts to argue from.