Anthropic Adds Watermarks In Claude Content As Deepfakes Run Amok

As AI-generated content floods the internet, differentiating machine output from human work has emerged as a defining problem. While regulators are push AI companies to label AI-generated content the default, major model developers are moving towards establishing that distinction as well.
For instance, 190 AI developers, including Google, Anthropic, Meta and OpenAI, have signed the EU AI Act’s Code of Practice on Transparency of AI-Generated Content.
With rising instances of deepfake, Indian authorities are also moving on a parallel track. Earlier this year, MeitY has proposed amendments to the IT Rules mandating platforms to label deepfakes and AI-generated content as synthetically generated information. This would be one of the first binding labelling regimes for AI content outside the EU.
Now, AI giant Anthropic has committed to make watermarking as a default across every Claude product, from generated APIs to Claude Code as well as enterprise deployments, worldwide. Claude models launched in the European Union on or after August 2, 2026 will carry machine-readable marks at launch, while older models will gain support during a transition period.
To note, its competitor OpenAI had also built a watermarking infrastructure in the past but never shipped it in ChatGPT. However, Google has already introduced SynthID, which does watermark Gemini generated text, image and audio output.
For context, a machine-readable mark is a symbol or hidden data layer on an item or file that a computer scanner or software can automatically detect and read. These aren’t meant for direct human reading.
How Claude Will Mark AI-Generated Content
Anthropic intends to use two complementary techniques for this purpose — embedded text watermarks and signed provenance metadata for files.
The watermark is woven directly into whatever Claude generates, invisible to readers, without changing its meaning, quality or readability. These watermarks are attached to the content irrespective of wherever the content is being used.
For supported file types such as .svg, .png and .jpg, Claude will attach signed provenance metadata that follows the C2PA open standard, the same framework Google and OpenAI have adopted for images and similar techniques have been explored for AI videos.
If a mark is detected on processing, it indicates that the content may have been processed by Claude, and Anthropic says it will help third parties build detection tools.
The marks apply across Claude platform (API), Claude, Claude Code, Claude Cowork and Claude Tag, and through cloud partners AWS, Google Cloud and Microsoft Foundry, wherever Claude is offered worldwide.
Anthropic also acknowledges the limits of its own system:
- A detected mark is a signal, not conclusive proof, of Claude’s involvement, since content may have been edited, excerpted or combined with other material after generation.
- The absence of a mark does not mean content was human-written, because heavy editing, paraphrasing, translation, short passages and stripped metadata can all erase the signal.
Why Text Watermarking Has Been A Hard Sell
This commitment comes after years of hesitation across the industry. OpenAI built a text watermarking tool reportedly capable of detecting GPT-4 output with more than 99.9% accuracy, but chose not to release it in ChatGPT, citing concerns about false accusations as well as user backlash.
An earlier OpenAI classifier was quietly shut down after it was able to identify AI generated text only 26% of the time. Instead, it adopted SynthID for images and audio in 2026, with no text watermark in sight. Google has gone further on text. DeepMind’s SynthID has watermarked Gemini text since May 2024 by embedding statistically detectable patterns in the model’s token selection.
On Anthropic’s watermarking, Contrails AI (an enterprise AI safety startup) CEO Amitabh Kumar said that this is the bare minimum a big tech can do. He, however, flagged major gaps with the Claude approach.
In this, watermarks are positioned as a probabilistic signal rather than a technical specification.
Details on how marks survive digital copies remain unclear, the third parties who will get detection access are unnamed, and there is little sign of interoperability with SynthID or whatever Meta builds.
He added that tutorials on generating content on cloud platforms without watermarks already circulate on Reddit and YouTube.
“It’s just a drop in the ocean of the damage deepfakes are causing,” Kumar said.
In India, the damage is already visible. Deepfakes went mainstream during the general election, when AI-generated videos of politicians and celebrities went viral, and have since hardened into tools of fraud and misinformation.
Scammers are using such content to promote fake investment schemes, manipulate public opinion and extract money.
The problem is particularly challenging because deepfakes can appear convincing and spread rapidly across social media and messaging platforms.
While the government and platforms are tightening rules around AI-generated content, detection and enforcement remain difficult. The growing accessibility of generative AI could make deepfake-enabled fraud an even bigger challenge.
The question now is whether Claude’s rollout avoids the failures that have slowed OpenAI and Google. Watermarks survive only if detection tools are open, standards stay interoperable, and providers resist the temptation to treat marking as a compliance checkbox.
Anthropic has promised technical guidance on its marking and detection approach, and the industry will be watching how quickly that promise translates into marks that actually hold up.
Edited by Akshit Pushkarna
Creatives by Varshita Srivastava
The post Anthropic Adds Watermarks In Claude Content As Deepfakes Run Amok appeared first on Inc42 Media.


Superadmin 










