On Monday October 5, OpenAI announced textGrain, an invisible watermark that will be applied to text generated by ChatGPT and Codex for users in the European Union over the coming weeks, and made available as an opt-in flag for API customers worldwide, off by default. The move complies with the EU AI Act's transparency rules, which took effect on August 2 and require generative AI providers to make their output identifiable in a machine-readable form, as TechCrunch reported.
Most coverage framed this as a compliance story, and the Register's verdict (OpenAI rolls out weak sauce watermarking for AI text) was that the compliance bar itself is low. But the accompanying technical report, textGrain: Entropy-Calibrated Watermarking for Language Model Text, is more interesting than the politics. Read it as an engineer and the whole product snaps into focus: the watermark is not ink stamped on text. It is a purchase. OpenAI buys a statistical signal by spending a budgeted amount of the model's own sampling entropy, and every strength and every weakness that the press coverage argues about follows from one inequality in that trade.
What textGrain Actually Does
The report, dated October 5 and co-authored with researchers from the University of Pennsylvania and Yale, describes a decoding-time watermark. At each generation step, a secret key plus a window of preceding tokens produces pseudorandom values. The vocabulary is partitioned into keyed blocks, and the model's next-token distribution over those blocks is nudged by an optimal transport problem whose costs come from Gumbel random variables, subject to an entropy budget that caps how much randomness the nudge may consume. A token is then sampled inside the chosen block using the original relative probabilities.
Three properties are worth stating precisely, because they are the design promises:
- Unbiasedness on average. If you average the watermarked output distribution over randomly drawn keys, you recover the exact unwatermarked distribution. The model's headline capabilities, measured across many keys, are provably untouched.
- A priced budget. The KL divergence between the coupling and independence equals the mutual information between the generated token and the keyed randomness, which equals the average entropy removed. The budget parameter beta directly says: we are willing to sacrifice this fraction of next-token entropy in exchange for provenance.
- A self-contained detector. Detection needs only the text, the tokenizer, the configuration, and the key. It reconstructs the pseudorandom values, turns each scored token's cost into a score, sums them, and compares the total to its null distribution, which the report derives as Gamma(n, 1), where n is the number of scored positions. Watermark is declared above the one-minus-alpha quantile, with a target false-positive rate of 1 percent.
There is real systems work attached, too. Each transport solve costs O(|A| + N_it·Bm) operations, where |A| is the active vocabulary support, B the block count, m the column count, and N_it the solver iterations; solves across a batch of positions stack into shared matrix ops. Appendix B shows the watermark composed with speculative decoding under a shared key, draft model and target model both watermarked, with a proof that verification preserves the target's watermarked output law. That appendix matters commercially: without it, watermarking would fight the inference optimizations every serving stack runs.
The Numbers OpenAI Published
The report and blog state detection rates that secondary outlets relayed, all of them OpenAI's own measurements, not third-party results. At the 1 percent false-positive setting, as The Decoder summarized:
| Condition (400-token passages unless noted) | Reported detection rate |
|---|---|
| Psychology prose, 400 tokens | about 95 percent |
| Psychology prose, 200 tokens | about 80 percent |
| Math content, 400 tokens | substantially lower, roughly 60 percent |
| 400 tokens, 10 percent of words swapped for synonyms | about 66 percent (down from about 92) |
| 400 tokens, 25 percent of words swapped | about 17 percent |
On output quality, OpenAI claims no significant differences with the watermark on or off across eight benchmarks including GPQA Diamond, BrowseComp, and DeepSWE, tested on its frontier model Astra. That is a vendor claim, and it is also narrower than it sounds: benchmark scores measure whether the right answer survives, not whether prose rhythm or word choice changed. The product-level rules round it out. Watermarking rolls out to all paid and free ChatGPT and Codex plans in the EU, API watermarking is opt-in worldwide including through cloud partners such as Azure, the detector is gated to approved researchers and expert organizations via an application process under the EU's Code of Practice, and OpenAI says it plans to open-source the technology.
Our Read: Every Weakness Is the Same Inequality
Here is the data-scientist's angle that the report itself supplies, almost as a footnote. Theorem A.4 bounds the signal: the mutual information between token and key, the entire detectable footprint of the watermark, is at most min{H(ρ), log m}, the entropy of the block distribution or the log of the column count, whichever is smaller. And the entropy budget only lets you buy up to beta times H(P) of it. Detection power is then a standard hypothesis-testing fact: the evidence from n scored tokens accumulates roughly linearly in n, while the noise in the Gamma(n, 1) null grows like sqrt(n). Power needs the per-token mean shift to clear noise that shrinks only as 1/sqrt(n).
Once you see that, all four published weaknesses are one story, not four.
Short text fails (80 percent at 200 tokens versus 95 percent at 400) because n is small, so the sqrt(n) noise dominates the linear signal. Nothing about the watermark is different in a tweet; the statistics simply run out of samples.
Math and code fail hardest (about 60 percent) for the most interesting reason: the budget is a fraction of H(P), the entropy of the model's own next-token distribution. In a math derivation or a code listing, the model is confident. H(P) is tiny. A tiny fraction of a tiny number buys almost no signal, exactly where the model has no freedom to pick a different synonym without making the output wrong. The Register noticed the practical consequence for functional text; the report's math says it is not a bug to be tuned away. Low-entropy regimes are information-theoretically cheap to watermark only if you are willing to distort the answer, which is the one thing the no-quality-loss claim rules out. For software engineers this lands awkwardly: Codex output is on the EU default list, yet generated code is the genre where the detector is weakest, and provenance for machine-written code is arguably where the Act's transparency goal has the most teeth.
Editing strips the mark (92 percent to 66 percent at 10 percent synonym swaps, down to 17 percent at 25 percent) because a swapped token is, from the detector's viewpoint, a token whose relationship to the key was destroyed. The report is candid that scoring only counts the first occurrence of each distinct context window and that the evidence is a sum over scored positions. Every edited token is one term deleted from that sum and, when the edit shifts context, neighbors can lose their scored alignment too. This matches everything known about generation-time watermarks since the paraphrasing-evades-detectors literature; textGrain did not escape it, it priced it.
Unbiasedness is the marketing shield, and it is real but narrow. Averaging over keys preserves the distribution, and the report is careful to say the entropy identity holds on average across keys, not per key, and that the numerical solver's achieved loss can drift above or below the requested budget. Fine print an engineer should respect: a strict budget claim, in the report's own words, requires checking the achieved loss separately.
There is a market read in the same math. Anthropic applies Google's SynthID-Text watermark globally and by default; OpenAI defaults it on only in the EU and leaves the API opt-in everywhere else. TechCrunch notes that OpenAI built an earlier text watermark and sat on it, partly for fear users would leave for unmarked rivals, and that Anthropic's global rollout drew user backlash. When the watermark costs entropy, shifts word choices, and is removable with a quarter of the words swapped anyway, the rational product posture is: comply exactly where the law says, sell the option elsewhere, keep the exit cost for your own customers at zero. The weak detection numbers make that posture cheaper to hold, not more expensive.
The queue-design echo from our piece on Google's bug bounty pause is worth naming too. A watermark detector is a classifier whose whole value depends on base rates and error asymmetry. OpenAI itself warns that a missing watermark does not prove human authorship: the text could be short, edited, translated, or from another lab's model. At 1 percent false positives and 60 to 95 percent true positives depending on genre, no institution should attach consequences to a detector verdict yet, which is precisely why the detector is gated to approved researchers. The compliance deliverable is a machine-readable mark, not a machine-readable judgment.
Outlook
Three things to watch. First, whether OpenAI ships the promised open-source release; if third parties can run detectors or attack them, the vendor-published rates above stop being the only numbers anyone has. Second, whether the EU-only default holds once API customers see watermarking as a trust feature rather than a tax; the opt-in flag is a pricing experiment wearing a compliance hat. Third, whether the quality claim survives scrutiny beyond benchmark scores: the Astra results on GPQA and friends show the right answer survives the nudge, but nobody outside OpenAI has measured whether the sentences themselves feel different.
For now the honest summary is short. OpenAI built a watermark that is statistically clean, systems-compatible, legally sufficient, and, by its own published numbers, easy to launder with light editing. That combination is not an accident of engineering. It is what you get when the law demands provenance at near-zero quality cost, and entropy is the only currency available.