Simon Willison·· 2 d ago
Quoting Anthropic Frontier Red Team
Quoting Anthropic Frontier Red Team
AI summary
Anthropic Frontier Red Team evaluated several models on 100 random tasks from its internal Binary Exploitation benchmark. GLM-5.3 achieved full control-flow hijacking in 4% of trials, versus 6% for Claude Mythos Preview. The report says earlier models such as Claude Opus 4.6 and GLM-5.2 had not succeeded on these tasks, suggesting a meaningful capability threshold has been crossed.
Selection record
Threshold 76Media and individualsFirst 62Second 68
Not admittedSum of both 130 < twice the threshold 152
- Source tier
- Media and individuals; this tier's threshold is 76
- Pre-filter
- passed:引用Anthropic红队评测AI模型网络攻击能力
- Same event
A model scores each item twice, independently, against one written standard, out of 100. An item is admitted only when the two scores add up to twice the threshold. The threshold is set per source tier.
Source: Simon Willison · simonwillison.net