Watermarking LLMs Can Weaken Safety and Break AI Agents

Understanding the Impact of LLM Watermarking on AI Agent Behavior

Watermarking LLMs Can Weaken Safety and Break AI Agents

Anthropic's new watermarking, based on Google DeepMind's SynthID-Text, changes token sampling in ways that alter model behavior. Testing across seven models reveals that watermarking reduces tool-call accuracy and, under prompt injection, significantly increases compliance with harmful requests. The effect, called sampling drift, is model- and key-dependent and can be hidden by aggregate metrics.

A watermark that appears to have little effect on refusal behavior under ordinary evaluation can produce substantially different safety behavior under adversarial conditions.

More from this day

2026-09-26