- OpenAI will add invisible watermarks to text responses from supported versions of ChatGPT and Codex in the European Union over the next few weeks.
- The company developed the technology in response to EU AI Act requirements for machine-readable identification of generated content.
- API customers worldwide will be able to enable watermarking voluntarily for specific models, but the feature will remain disabled by default.
- OpenAI said editing, translation and limited word-choice flexibility can reduce detection rates.
OpenAI will introduce invisible watermarking for text responses from supported versions of ChatGPT and Codex in the European Union over the next few weeks. The technology is intended to help determine whether text was created or processed by OpenAI systems as generative AI providers face EU requirements to make generated content identifiable in machine-readable form.
OpenAI introduced the approach in response to the EU AI Act. The feature will also become available globally to API customers, who will be able to enable watermarking voluntarily for specific models. It will remain disabled by default.
How textGrain works
The system, called textGrain, places an invisible statistical signal in a model’s word choices. A dedicated detector then analyzes the text to determine whether it contains that signal.
OpenAI said textGrain performed as well as or better than other approaches it tested, including SynthID for text. The company cautioned, however, that laboratory results do not guarantee the same effectiveness under real-world conditions.
At a 1% false-positive rate, the detector found the watermark in about 80% of 200-token excerpts and about 95% of 400-token excerpts on topics such as psychology, according to OpenAI. Detection rates were lower for mathematics because models had less freedom in choosing words.
Editing also significantly affected performance. Replacing 10% of words with synonyms reduced the detection rate from about 92% to 66%, while replacing 25% of words lowered it to 17%.
Watermarks provide limited provenance information
OpenAI said a watermark provides only a limited signal about a text’s provenance. A detection result does not reveal how much of the text a human contributed, who owns or is responsible for it, who used the OpenAI system, which prompt or conversation produced it, or whether the information is true or false.
The absence of a watermark also does not establish that a person wrote the text. OpenAI said detection could fail because a passage is too short, has been edited or translated, came from a different model, or was created before the technology was introduced.
Detector access will initially be restricted
OpenAI has opened applications for approved researchers and expert organizations seeking access to the detector. The company cited the risks of false-positive and false-negative results and said it would initially provide access individually to evaluate the system’s reliability and responsible uses.
The company also plans to open-source textGrain so other developers can build their own tools using the technology. OpenAI said testing of the new advanced Astra model found no significant difference in benchmark results with and without watermarks.
OpenAI said no single method can fully verify content provenance. It plans to continue improving watermark detection, including by assessing its resistance to editing and translation, and eventually aims to distinguish more effectively between AI-assisted writing and full AI authorship.
Separately, the source previously reported that OpenAI was seeking to raise $30 billion at a $1.4 trillion valuation.
Source: Incrypted
