AI Model Watermarking Changes Agent Behavior
Lasso Security found that SynthID-Text watermarks, which the EU AI Act requires in AI content, alter tool calls and model refusals. On the BFCL v4 benchmark, accuracy dropped in six of seven models, and prompt injection attack success rates increased.
- Watermarks reduced tool call accuracy in 6 of 7 models on BFCL v4
- Prompt injection attack success notably rose on watermarked models
- Effect also impacts third-party agents accessing Anthropic API
- Lasso urges including watermarks in agent security evaluation
Read next
AI