
OpenAI has introduced a new framework for tracking, investigating, and disclosing model misalignment. The company has also published six reports covering unexpected model behavior observed during training and evaluation over the past six months. Continue reading “OpenAI introduces framework for reporting model misalignment, publishes six reports”
