OpenAI News·· 15 d ago
Our framework for reporting model misalignment
Our framework for reporting model misalignment
AI summary
OpenAI released a framework for tracking, investigating and disclosing model misalignment, alongside six reports of unexpected or concerning model behaviour.
Selection record
Threshold 60Official, first-handFirst 62Second 58
AdmittedSum of both 120 ≥ twice the threshold 120
- Source tier
- Official, first-hand; this tier's threshold is 60
- Pre-filter
- passed:OpenAI模型失准报告框架,实质AI安全
- Why it was chosen
- The model misalignment tracking and disclosure framework, with six abnormal-behaviour reports, explains OpenAI’s safety disclosure process.
A model scores each item twice, independently, against one written standard, out of 100. An item is admitted only when the two scores add up to twice the threshold. The threshold is set per source tier.
Source: OpenAI News · openai.com