Watermarking AI-generated text: effects for the public and businesses

Views: 34

The EU AI Act requires GPAI providers to comply with the copyright rules, respect rights for text and data mining, and publish a summary of the content used to train their models; but it does not create new or replace existing intellectual property and copyright laws. In the standard GPAI, the AI providers that generate or manipulate text, images, audio, video, etc. must make the output technically identifiable as AI-generated or manipulated, e.g. through machine-readable metadata.  

Background
Basically, the text watermarking is a technique for embedding hidden information within textual content to verify its authenticity, origin or ownership. With the rise of generative AI systems using large language models, there has been significant development focused on watermarking AI-generated text.
The AI text watermark represents a kind of “hidden statistical pattern” woven directly into machine-generated writing. Major AI providers, such as Anthropic for Claude and Google for Gemini, use these embedded markers specifically to comply with transparency laws like those already enforced by the EU AI legislation.
Watermarking does not necessarily mean adding a visible label for the person viewing the content. Simple editing tools that do not substantially alter the input data or its meaning are exempt.
Additionally, on watermarking in: https://www.layer3labs.io/guides/ai-watermarking

For transparency risk, the EU AI Law (transparency obligations in art. 50 took effect in August 2026) focuses on AI models that could make it difficult for people to distinguish between the AI-generated or human-created content. The concern is mainly deception, impersonation, misinformation, manipulation and consumer fraud. Such AI models/systems are generally allowed, but providers and deployers must make the use of AI sufficiently transparent. The rule has extraterritorial reach: it applies to US and UK companies whose AI output is used in the EU.
More about art.50 in: https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-50

Practical embedment
AI watermarking embeds an invisible, machine-readable signature directly into generated text, images and/or audio. It subtly alters statistical patterns or pixel data during creation; this mark remains hidden to humans but can be detected by specialized scanning algorithms to verify authenticity.
Thus, the AI watermark can detect the AI-generated content: hence, companies have embedded invisible statistical patterns or digital fingerprints directly into text, images or video the moment the AI creates them. Computers use a secret key or rules to spot these hidden signals.
= In media and plain text processing, the specialized tools and platforms can read these hidden data layers to verify if a file came from an AI model.
= Some major providers—such as Anthropic’s Claude—intentionally incorporate statistical patterns or structured token choices into their text outputs to comply with transparency envisioned by the EU AI Act’s regulations.
= Other models like OpenAI’s have sometimes leaked hidden Unicode space characters (like zero-width spaces) into text, though developers state these are unintended quirks of reinforcement learning rather than deliberate tracking watermarks.
= In the statistical bias, the AI text inherently relies on predictable patterns and word-choice probabilities, which automated detectors use as a behavioral “watermark.
There are some AI watermarks vital details: thus, in the text watermarks, the AI slightly changes its word choices using a hidden mathematical pattern. The text reads normally to humans, but a special detector counts the favored words to prove that the AI made it.
= In image and video watermarks: some AI tools/models like Google DeepMind SynthID embed invisible signals into pixels or audio frames that survive cropping, filtering or compression.
= In compliance: the EU AI Act regulations require digital developers to mark synthetic outputs in machine-readable formats.
Anthropic intends to “watermark texts” generated by its models, including Claude, to comply with the EU digital regulations; the AI model-maker confirmed the watermarking in an updated support page. EU AI Act’s Transparency Code, which took effect on August 2, requires AI companies to mark AI-generated or edited content in a way other systems can identify them.
Thus, AI providers must: a) design AI systems in a way that ensures that customers are explicitly informed whenever they interact with an AI system directly; and b) add machine-readable marks to enable the detection of AI-generated or manipulated content.
More in: https://digital-strategy.ec.europa.eu/en/policies/guidelines-ai-transparency-obligations

The EU AI legislation for business
Some digital experts, like Romain Digneaux, Public Policy Manager at Proton, already explained how the EU AI legislation affects those who use AI-powered products or services in the EU, whether for personal or business purposes.
One effect is already visible in the labels or disclosures on platforms such as Instagram and TikTok indicating that content was generated or manipulated using AI, which became applicable in August 2026 under the transparency rules.
The EU-wide AI legislation has been implemented in stages – more on that below; the latest one (from August 2026) is about the ways the AI act treats various risks: e.g. from unacceptable-risk and high-risk systems to transparency-risk and low-risk ones.
More on AI Act in: https://ai-act-service-desk.ec.europa.eu/en/ai-act-service-desk

Thus, during the turning of legal obligations (attached to GPAI models), the corporate community has to validate how the AI law restricts the use of people’s data for training, for protecting intellectual property, etc. Specifically, for businesses, compliance begins with establishing the company’s role and the kind of AI it uses.
For example, in the Claude-generated text, watermarking doesn’t change the meaning or experience for the person reading it, but if someone wanted to check whether the text was likely generated by Claude, the watermark allows to do so; and the watermark only applies to words Claude chooses.
When Claude proofreads text written by a person, what it gives back has generally only been lightly edited; because nearly all the words are the person’s, there’s very little (if anything) for the watermark to attach to. Depending on the length of the text and how heavily Claude has edited it, those changes might not be enough to make Claude’s involvement detectable. The more Claude writes, the more decisions it has to make, and the more space there is for a watermark.
But the watermark cannot be traced back to an organization; the watermarking applies to Claude and its outputs, and it doesn’t identify anything to do with individual users. There’s nothing in the watermark, or its key, that would allow anyone to recover any information about the user, their organization, or their chats with Claude.
Basically, the watermarking in the Claude’s outputs have been implemented in view of complying with the EU AI Act. Anthropic, along with several other major AI model providers and around 190 total signatories, signed the EU Code of Practice on Transparency of AI-Generated Content in July 2026. This requires AI system providers to use methods of “marking” AI-generated text; hence, the Anthropic is applying watermarking globally, including the EU.
https://www.anthropic.com/news/claude-text-watermark

However, it is not illegal to publish a book/article written or assisted by AI, but it comes with major risks: the user may face charges concerning the copyright ownership, platform rules and accidental plagiarism and similarity issues.
https://www.double9books.com/blogs/blog/is-it-legal-to-publish-an-ai-generated-book

EU AI’s final timetable
= From the end of 2026: The ninth prohibition covering AI-generated CSAM and non-consensual sexually explicit or intimate depictions takes effect. Providers of AI systems placed on the market before 2 August 2026 that generate synthetic text, images, audio and/or video must also comply with the Act’s machine-readable marking requirements by this date.
= Since 2 August 2027: the EU member states must have at least one national AI regulatory sandbox available by this date for companies to develop and test AI systems under regulatory supervision.
= From the end of 2027: high-risk AI rules begin to apply to the digital systems used in sensitive areas such as biometrics, critical infrastructure, education, employment, essential services, law enforcement, migration and the administration of justice.
= From August 2028: high-risk requirements begin to apply to AI systems/models that are part of, or safety components of, regulated products covered by the EU-wide product legislation, such as certain machinery, medical devices, toys, lifts, etc.
Source: https://proton.me/blog/eu-ai-act#timeline

One thought on “Watermarking AI-generated text: effects for the public and businesses

  1. The explanation of AI watermarking is very informative, especially the discussion of transparency, machine-readable identification, and the EU AI Act. The impact on businesses and the importance of responsible AI use are particularly interesting.

Leave a Reply

Your email address will not be published. Required fields are marked *

19 + 18 =