Anthropic: Zhipu GLM-5.3's Cyberattack Capabilities Approach Claude, but Safety Guardrails Remain Weak
nashnova research
Anthropic reports that Zhipu's latest model GLM-5.3 nearly matches its own Claude in exploit-development benchmarks, yet its safety mechanisms are trivially easy to strip away — offensive capability is approaching the frontier while defenses lag far behind.
How strong is GLM-5.3's attack capability?
Out of 410 attempts, GLM-5.3 produced working end-to-end exploit code 50 times; Claude Mythos Preview scored 56 — a narrow gap.
In the harder binary-exploitation test — hijacking a program's control flow — GLM-5.3 succeeded 4% of the time versus Claude's 6%; Kimi K3 and DeepSeek V4.1-Flash both scored 0%.
This means → on the question of "can this model write a real, usable attack," GLM-5.3 is already the strongest open-weight model from China and is closing in on the U.S. frontier.
Why are safety guardrails practically useless?
Anthropic found that after modifying GLM-5.3's weights and removing its safety layers, the model's refusal rate dropped from above 90% to roughly 2%–12%.
In plain terms = the model used to reject nine out of ten malicious requests; strip the guardrails and it refuses almost nothing.
This reflects a fundamental tension in open-weight models — models anyone can download and modify: when capability is open, the safety mechanism is exposed to everyone, including bad actors.
What do officials and Zhipu say?
The U.S. Commerce Department's CAISI (Center for AI Standards and Innovation) assessed in September that GLM-5.3 is the most cyber-capable open-weight model released to date, trailing the U.S. frontier by roughly four months.
Zhipu's head of global affairs, Zixuan Li, responded that the model is also used for cyber defense — it has helped protect 389 open-source projects and identified 4,249 potential vulnerabilities.
This means → the same capability cuts both ways; the debate is not about the technology itself but about who gets the tool and whether it can be effectively constrained.
What comes next?
Anthropic's report is the first to present Chinese open-weight models' offensive cyber capability in hard numbers, moving the conversation from qualitative concern to quantified reality.
The key follow-up: whether this data pushes regulators toward stricter safety reviews of open-source AI models — especially given that freely downloadable weights may render existing safety frameworks inadequate.
In plain terms = "open models could be misused" was a worry before; now Anthropic has put specific numbers on the table — the question is no longer "will it happen" but "how soon."
市场有风险,内容仅供研究参考,不构成投资建议。
