Skip to content
Simon Willison·· 2 d ago

Quoting Anthropic Frontier Red Team

Quoting Anthropic Frontier Red Team

AI summary

Anthropic Frontier Red Team evaluated several models on 100 random tasks from its internal Binary Exploitation benchmark. GLM-5.3 achieved full control-flow hijacking in 4% of trials, versus 6% for Claude Mythos Preview. The report says earlier models such as Claude Opus 4.6 and GLM-5.2 had not succeeded on these tasks, suggesting a meaningful capability threshold has been crossed.

Selection record

Not admittedSum of both 130 < twice the threshold 152

Source tier
Media and individuals; this tier's threshold is 76
Pre-filter
passed:引用Anthropic红队评测AI模型网络攻击能力
Same event

A model scores each item twice, independently, against one written standard, out of 100. An item is admitted only when the two scores add up to twice the threshold. The threshold is set per source tier.

Source: Simon Willison · simonwillison.net