Watermarking LLMs Can Weaken Safety and Break AI Agents
Understanding the Impact of LLM Watermarking on AI Agent Behavior

Anthropic's new watermarking, based on Google DeepMind's SynthID-Text, changes token sampling in ways that alter model behavior. Testing across seven models reveals that watermarking reduces tool-call accuracy and, under prompt injection, significantly increases compliance with harmful requests. The effect, called sampling drift, is model- and key-dependent and can be hidden by aggregate metrics.
A watermark that appears to have little effect on refusal behavior under ordinary evaluation can produce substantially different safety behavior under adversarial conditions.