AI Security Research Roundup: Multi-Scanner Guardrails Bypassed, Reward Models Stolen
New AI security research from October 4–11, 2026 shows the BRANCH attack bypassing six multi-scanner guardrail systems with 100% success, while the Reward Stealing Attack recovers an aligned LLM's reward model black-box. Defenses ASPIRE, LADE, AdaGuard and AttestMCP were also presented.
- BRANCH hit 100% attack success across 6 guardrail systems in 120 scenarios with 72% fewer queries
- Reward Stealing Attack recovers the reward model via inverse RL and elicits harmful outputs without white-box access
- AttestMCP cut attack success on MCP agents from 53.7% to 12.4%
- Anthropic's open-source scanner found over 29,000 bugs
Read next
Security