Anthropic: GLM-5.3 Safeguards Bypassed in 64–100% of Tests
Anthropic published a report claiming that the guardrails of GLM-5.3, an open-weight model from China's Z.ai, can be bypassed with simple techniques in 64% to 100% of simulated tests. Researchers used abliteration to strip the model's refusals, producing a version that scored 3%, 2% and 12% on JailbreakBench, HarmBench and StrongREJECT versus about 90% for the original.
- Anthropic bypassed GLM-5.3 safeguards in 64–100% of simulated tests
- Abliteration cut the model's scores to 3%, 2% and 12% on three benchmarks
- The original GLM-5.3 scored about 90% on JailbreakBench, HarmBench and StrongREJECT
- BBC: Moonshot probes Kimi models after jailbreak reports on bioweapons guidance
Read next
AI