On October 5, OpenAI explained how it plans to meet the EU AI Act provision requiring machine-generated text to be identifiable in a machine-readable way. The method is called textGrain, and it puts an invisible statistical signal into the model’s word choices. A detector then looks for that signal.
The rollout comes in three parts. From today, API customers worldwide can switch the watermark on for selected models — in the API it stays off by default. Over the coming weeks, ChatGPT and Codex output in the EU gets the mark, with no opt-in there. And the detector isn’t being published: you apply for it, and OpenAI is handing it only to approved researchers and expert organizations at first.
For images and audio nothing changes. openai.com/verify and the Content Provenance API stay open to everyone.
The numbers OpenAI puts on the table itself
At a target false positive rate of one percent, the detector finds the mark in roughly 80 percent of 200-token passages and roughly 95 percent of 400-token passages — measured on content such as psychology. For mathematics the rate drops substantially, because word choice there leaves so little room. The shorter and the more formulaic the text, the worse it gets.
Editing makes it starker. In 400-token passages, detection falls from about 92 to 66 percent when ten percent of the words are swapped for synonyms. At 25 percent it’s down to 17.
What the mark does not do gets its own section in OpenAI’s post: it doesn’t measure human contribution, it doesn’t settle ownership or responsibility, and it doesn’t identify a person.
Quality seems untouched. OpenAI shows eight benchmarks for Astra at max effort, watermarked and not; the numbers sit close together throughout — GPQA Diamond at 93.94 against 94.44 percent, and Terminal-Bench 4.0 at 56.06 against 53.90, this time in favour of the watermarked run.
Anthropic did this in August
There’s a reason the approach sounds familiar. Anthropic has been watermarking Claude’s text output since August, and I pulled apart how it works back then. Same idea, same weak spot: rewrite the text and most of the mark goes with it.
OpenAI says its own tests put textGrain level with or ahead of Google’s SynthID for text. The technology is meant to be open-sourced later.
Who gets marked depends on your postcode
Two sentences in the same announcement sit in a striking relationship. In the EU, ChatGPT text gets the mark without anyone being asked. In the API it’s off everywhere until a developer turns it on — which is exactly where text gets produced in bulk and passed along.
This isn’t a dodge around the AI Act; that provision aims at the provider’s relationship with end users. It just shows how far the obligation and the actual problem sit apart. And the detector, the one thing that would make any of it checkable, is behind an application form for now.