OpenAI announced text watermarking in the EU on 5 October 2026, and the announcement splits into three separate things. Per OpenAI's EU text provenance announcement:
- API customers worldwide can opt in to text watermarking for select models. It stays off by default.
- Over the coming weeks, OpenAI will add an invisible watermark to eligible ChatGPT and Codex text output in the EU, across all plans. It isn't a global default at launch.
- A detector exists, but access is initially limited to approved researchers and expert organizations.
The driver is Article 50(2) of the EU AI Act, which requires providers of systems that generate synthetic text to mark outputs in a machine-readable, detectable way. The coverage we saw retells the announcement. The parts that matter for an engineering team are in the details: where the opt-in lives, what the watermark can't prove, and what OpenAI hasn't said about code.
What OpenAI is changing (and where)
Three scopes, three different audiences.
The API opt-in is global, not EU-only. It's a settings toggle rather than a request parameter. OpenAI's help center article on provenance signals says to turn on "Allow text watermarking," select your models, and save. You set the organization default under Organization settings > Data controls > Text provenance, and you can override it per project under Project Settings > Text provenance. You don't add a marking to each response.
ChatGPT and Codex are EU-only for now, and the rollout is "over the coming weeks," so don't assume it's live for your users today.
The detector is the third scope, and it isn't yours. OpenAI says the tool reports whether an OpenAI watermark is detected without identifying the user or revealing prompts. Applications are open, but turning on API watermarking doesn't get you access.
OpenAI also says it's working with cloud partners to offer watermarking for OpenAI model outputs "in the coming weeks," and that availability "may vary by output and by partner." If you reach OpenAI models through AWS, our look at OpenAI models reaching AWS teams through Bedrock is the context. We found nothing on what cloud availability will look like beyond that sentence, so we're not guessing.
What the watermark is and what it can't tell you
OpenAI calls the technique textGrain. It adds an invisible statistical signal to the model's word choices. There are no hidden characters, invisible spaces or unusual punctuation, so copy-pasting adds no hidden material. The signal lives in the wording itself. OpenAI says that if you preserve the wording, the signal is expected to remain, and that rewriting, paraphrase or translation make detection less reliable.
How much less? OpenAI's figures, at a 1% target false positive rate:
- About 80% of 200-token passages are detected, and about 95% of 400-token passages (psychology content). Mathematics is lower.
- On 400-token passages, replacing 10% of words with synonyms cuts detection from about 92% to 66%. Replacing 25% cuts it to 17%.
- Across 24 EU languages, the help center reports Spanish highest at 69.0% and Romanian lowest at 42.2%. That test used 500 synthetic prompts translated into each language, so don't over-read it.
So this is a probabilistic signal that light editing erodes. OpenAI is explicit about the limits: a watermark doesn't measure human contribution, establish ownership or responsibility, identify the user or verify accuracy. And no watermark doesn't prove a person wrote the text.
On quality, OpenAI reports no meaningful differences on its benchmarks. One example it gives is 72.80% unwatermarked against 71.68% watermarked on a coding benchmark. That's OpenAI's own testing, so run your own evals before you assume it.
What it means for Codex and code
OpenAI's announcement says "eligible ChatGPT and Codex text output" will carry the watermark. It doesn't define "eligible," and it doesn't say whether code from Codex is watermarked. We won't claim either way.
What OpenAI does say: code is harder to watermark because there are fewer plausible choices for what comes next than in ordinary prose. It also says the Code of Practice "does not require watermarks in outputs shorter than 200 tokens ... or in code snippets," and that short passages usually don't contain enough text for reliable detection.
The legal side is narrower than people assume. The Commission's FAQ on Article 50 transparency obligations lists source code outputs among the outputs exempt from marking where they're "intended to be exclusively communicated from machine-to-machine and processed automatically without any exposure to humans." Our reading: code a developer reads in a pull request has human exposure, so don't read this as "code is exempt."
Our inference, not OpenAI's statement: the prose around code (PR descriptions, commit messages, docs, comments, and customer emails drafted in ChatGPT) looks like ordinary text, so it's the likeliest to carry a signal. Whether a given Codex surface does is something OpenAI hasn't documented.
What Article 50 requires of the provider, and who that is
Article 50(2) of Regulation (EU) 2024/1689 on EUR-Lex (as originally adopted) says providers of AI systems generating synthetic text must ensure outputs are marked in a machine-readable format and detectable as artificially generated. The solutions have to be effective, interoperable, sturdy and reliable as far as technically feasible. The marking duty doesn't apply where a system only performs an assistive function for standard editing or doesn't substantially alter the input or its semantics.
Who's the provider? For ChatGPT, OpenAI is the obvious candidate as the operator. If you ship your own product under your name on top of an API, you may hold that role yourself for that product, and a contractor building it complicates the answer. That's a call for your counsel.
Breaching the transparency obligations carries fines up to EUR 15,000,000 or, for an undertaking, up to 3% of total worldwide annual turnover, whichever is higher (Article 99(4)(g), same regulation).
On dates, the Commission FAQ says Article 50 applies from 2 August 2026, with a limited grace period to 2 December 2026 that the FAQ describes as "envisaged only for AI systems placed on the market before 2 August 2026 and only as regards the marking and detection obligation." Content generated before 2 August doesn't need retroactive labels. A system launched after that date doesn't get the transition. The other duties in Article 50, including the deployer duty on public-interest text under Article 50(4), aren't met by provider watermarking. Our guide to the four Article 50 duties covers them, so we won't repeat it here.
Should you turn on API watermarking?
The mechanics are cheap. It's off by default, you set it per organization or project, you pick models, and there's no per-request work. The model list isn't published as a list. You see which models are available for watermarking in your settings, and OpenAI says coverage extends to all legacy models over the coming weeks. The API changelog we read on 7 October had no watermarking entry, and we found no API parameter documented, so don't build on one.
Our judgment: turn it on if you ship generated text to EU users under your own name and you have no other marking in place. The quality cost looks small on OpenAI's numbers.
Two cautions. First, check with counsel whether OpenAI's watermark alone meets your own Article 50(2) duty. OpenAI says it can't advise on your legal obligations, and that watermarks and C2PA metadata don't replace visible labels, banners or other notices that may be required. Second, know the gaps: translation and rewriting weaken the signal, and enabling it gives you no way to check your own output, because detector access is separate.
What this means for your engineering team
This is a risk-posture decision, and four things are worth doing this month.
- Name an owner for the provider and deployer mapping under the AI Act, per product.
- Set the organization and project toggles deliberately instead of leaving the default by omission.
- Update your vendor questionnaire answers: which vendor marks which output, and that the detector is limited to approved researchers and expert organizations, so customers can't run it by default.
- Tell staff that eligible EU ChatGPT and Codex text is due to carry a signal that isn't proof of authorship, so no policy should treat a detection result as evidence about who wrote something.
If the provider question for software a contractor built is the open one, start with who's the provider when an outsourced team builds the AI feature. For other vendor changes with dates attached, see AI model releases and pricing: what changes for engineering teams.
FAQ
Does OpenAI watermark API output by default?
No. OpenAI says text watermarking stays off by default in the API. You opt in per organization or project in settings, and you choose which models.
Does the watermark apply to code from Codex?
OpenAI doesn't say. It says "eligible" ChatGPT and Codex text output is covered, doesn't define eligible, and states that code is harder to watermark. It also says the Code of Practice doesn't require watermarks in code snippets or outputs under 200 tokens.
Can I check whether text has a watermark?
Not yet, unless you're an approved researcher or expert organization. OpenAI says detector access is initially limited to that group, and enabling API watermarking doesn't change it.
Does the watermark replace an AI-generated label on my content?
No, according to OpenAI. It says watermarks and C2PA metadata don't replace visible labels, banners or other required notices. Deployer labelling under Article 50(4) is a separate duty.
When does Article 50 apply?
From 2 August 2026. Per the Commission FAQ, there's a transition to 2 December 2026 only for the marking and detection duty, and only for AI systems already on the market before 2 August.
