chiprook
← AI
AIOctober 1, 2026, 19:00

Anthropic: GLM-5.3 Safeguards Bypassed in 64–100% of Tests

Anthropic published a report claiming that the guardrails of GLM-5.3, an open-weight model from China's Z.ai, can be bypassed with simple techniques in 64% to 100% of simulated tests. Researchers used abliteration to strip the model's refusals, producing a version that scored 3%, 2% and 12% on JailbreakBench, HarmBench and StrongREJECT versus about 90% for the original.

Anthropic: GLM-5.3 Safeguards Bypassed in 64–100% of Tests
#Anthropic#Z.ai#GLM-5.3#Moonshot
Read next
AI

Z.ai and Concordia AI propose six stages for managing open-weight AI risk

AI

Aikido Security Releases Altar-1: Open-Weight Security Model Pruned From GLM-5.3 to 328 GB

AI

China's Z.ai disables AI coding assistant features after security issue

AI

GLM-5.3-Flash: 320B Open-Weight Model With 18B Active and 1M-Token Context