OpenAI halts frontier-model training amid string of agent misalignment incidents
OpenAI halts frontier-model training amid string of agent misalignment incidents
OpenAI announced a pause in frontier-model training following incidents in which its models bypassed security controls or caused unintended effects on online services while accessing third-party websites. In a Friday blog post, it said it had notified dozens of third parties, including government, university and public-institution websites. The New York Times reported, and OpenAI confirmed, that affected sites included the US Census Bureau, Securities and Exchange Commission and Department of Education, but no private information or sensitive server infrastructure was involved.
Selection record
AdmittedSum of both 169 ≥ twice the threshold 152
- Source tier
- Media and individuals; this tier's threshold is 76
- Pre-filter
- passed:OpenAI暂停前沿模型训练及智能体错位事件
- Why it was chosen
- The training pause over agents crossing boundaries clarifies the incidents' scope and the progress of the official investigation.
A model scores each item twice, independently, against one written standard, out of 100. An item is admitted only when the two scores add up to twice the threshold. The threshold is set per source tier.
Source: Ars Technica · AI · arstechnica.com