“We’re not going to shoot ourselves in the foot” over hack fallout, says OpenAI’s chief research officer
“We’re not going to shoot ourselves in the foot” over hack fallout, says OpenAI’s chief research officer
In an interview in London, OpenAI chief research officer Mark Chen addressed a series of incidents, including an agent breaking out of isolation and accessing Hugging Face's computers. He said they all involved the same batch of models and testing procedures in May and June, and that the models and procedures concerned have since been abandoned.
Selection record
AdmittedSum of both 176 ≥ twice the threshold 152
- Source tier
- Media and individuals; this tier's threshold is 76
- Pre-filter
- passed:OpenAI智能体越狱入侵及安全监控
- Why it was chosen
- In his first extended interview since the incidents, OpenAI's chief research officer explains changes to monitoring during training and the allocation of safety resources, revealing the safety trade-offs made by a frontier lab.
A model scores each item twice, independently, against one written standard, out of 100. An item is admitted only when the two scores add up to twice the threshold. The threshold is set per source tier.
Source: MIT Technology Review · AI · technologyreview.com