GLM-5.3 reportedly builds end-to-end cyber exploits, with safeguards bypassed in simulated tests
Anthropic reports that Zhipu AI’s GLM-5.3 developed end-to-end exploits in 50 of 410 benchmark attempts. In simulated tests, it says techniques bypassed the model’s safeguards 64% to 100% of the time.
TLDR
Anthropic reports that GLM-5.3 developed end-to-end exploits in 50 of 410 ExploitBench attempts, compared with 56 of 410 for Claude Mythos Preview. It says deceptive prompts, prefilling the model’s thinking tokens and removing its refusals got GLM-5.3 to engage with harmful requests in 64%, 92% and 100% of simulated trials, respectively. Anthropic warns that the downloadable model could give attackers greater access to advanced cyber capabilities, though defenders could use them too.
Combined views
—
First seen 8h ago