Import AI· Jack Clark·· 2026-07-27
Import AI 466: The bitter lesson for robotics, AIs complete week-long programming tasks; and OpenAI's accidental AI hacker
Import AI 466: The bitter lesson for robotics, AIs complete week-long programming tasks; and OpenAI's accidental AI hacker
AI summary
Epoch and METR released MirrorCode to test AI systems' ability to complete programming tasks requiring extended human effort. First announced in April, it is formally released after additional testing.
Selection record
Threshold 76Media and individualsFirst 71Second 71
Not admittedSum of both 142 < twice the threshold 152
- Source tier
- Media and individuals; this tier's threshold is 76
- Pre-filter
- passed:AI模型编程评测、机器人任务与安全事件
A model scores each item twice, independently, against one written standard, out of 100. An item is admitted only when the two scores add up to twice the threshold. The threshold is set per source tier.
Source: Import AI · importai.substack.com