Skip to content

#Anthropic

0 today

Sep 30

Wednesday
  1. Quoting Anthropic Frontier Red Team

    Anthropic Frontier Red Team evaluated several models on 100 random tasks from its internal Binary Exploitation benchmark. GLM-5.3 achieved full control-flow hijacking in 4% of trials, versus 6% for Claude Mythos Preview. The report says earlier models such as Claude Opus 4.6 and GLM-5.2 had not succeeded on these tasks, suggesting a meaningful capability threshold has been crossed.

Sep 22

Tuesday

Jul 27

Monday

Jun 22

Monday