Skip to content
OpenAI News·· 15 d ago

Our framework for reporting model misalignment

Our framework for reporting model misalignment

AI summary

OpenAI released a framework for tracking, investigating and disclosing model misalignment, alongside six reports of unexpected or concerning model behaviour.

Selection record

AdmittedSum of both 120 ≥ twice the threshold 120

Source tier
Official, first-hand; this tier's threshold is 60
Pre-filter
passed:OpenAI模型失准报告框架,实质AI安全
Why it was chosen
The model misalignment tracking and disclosure framework, with six abnormal-behaviour reports, explains OpenAI’s safety disclosure process.

A model scores each item twice, independently, against one written standard, out of 100. An item is admitted only when the two scores add up to twice the threshold. The threshold is set per source tier.

Source: OpenAI News · openai.com