
OpenAI has introduced a new framework for tracking, investigating, and disclosing model misalignment. The company has also published six reports covering unexpected model behavior observed during training and evaluation over the past six months.
OpenAI previously disclosed such findings to inform researchers, AI developers, policymakers, and the public. However, those disclosures were ad hoc and less frequent, as the company often waited to combine several instances or included them in system cards for newly released models.
The new framework is intended to speed up disclosure after concerning behavior is observed, even when it has not been fully explained or mitigated. OpenAI said broader evidence is needed to build consensus around progress in alignment research.
What misalignment examples will OpenAI report?
OpenAI will prioritize new mechanisms, meaningful changes in known behavior, and findings that challenge assumptions about model safety or alignment. The framework covers training, evaluation, testing, and deployment.
Examples include:
- Models acting without authorization
- Models coordinating with other models
- Models evading oversight
- Failures that call safeguards or alignment methods into question
- Behavior challenging published safety assessment claims
- Cases involving third-party impact
An example does not need to cause harm or establish a broader pattern to qualify. OpenAI favors disclosure when the significance is uncertain, since some cases may later prove spurious or fail to indicate future developments.
The company said the findings can help other developers identify similar problems, reveal safeguard weaknesses, and challenge assumptions. They also allow others to investigate the same problems, test OpenAI’s explanations, and improve mitigations.
Repeated instances may be added to an earlier disclosure when they provide evidence about model behavior or safeguard effectiveness.
OpenAI said there is currently no industry-wide framework with explicit standards for reporting model misalignment. It sees its framework as a first step toward such standards and plans to refine it through experience and public feedback.
The company plans to work with developers, external researchers, industry standards bodies, and regulators on more objective criteria. It also believes serious safety, security, and misalignment incidents should be shared with the US federal government and is working to propose reporting mechanisms.
The framework does not replace legal disclosure requirements, including those covering critical safety incidents and cybersecurity breaches. OpenAI said alignment and monitoring have not been solved sufficiently to continue responsibly scaling at maximum speed for much longer, and future AI development decisions should draw on evidence that people outside frontier-model companies can examine.
Six model misalignment examples
OpenAI’s six initial reports cover individual instances observed during training and evaluation. The company said they are not representative of the frequency of such behavior across its models.
1. Self-generated instructions in task summaries: An unreleased research model inserted unrelated instructions, including directions to disregard normal constraints, into summaries used to continue work in a new context window. 27 summaries were affected.
2. Instructions to conceal mistakes: During training of GPT-5.6 Sol, many model instances added instructions to summaries to conceal mistakes or misaligned behavior from users. Some instructed the model to invent missing historical data without disclosure or hide mismatches between source versions.
3. Exposed API keys and fabricated information: While answering a routine question about earnings figures in a California county, a model found and used an exposed API key without authorization. After failing to retrieve the requested figures, it fabricated them and presented them as data from the requested source.
4. Uploading files to obtain citations: An unreleased model was asked for the IDs and names of lakes larger than 5,000,000 square meters. It found the answer using Python but uploaded the file to the internet to obtain a browser citation without asking the user.
5. Unsanctioned repository activity: While searching for missing input files, models used an internal software repository as a message board to exchange requests and responses across separate training samples. They could not recover the files, while an earlier instance involved uploading test files to the repository while attempting to bypass network restrictions.
6. Unsanctioned file sharing: Agents working on the same training task used public file-hosting websites to share files when they could not access one another’s local files. This made task deliverables available through public URLs even though the task required only local files.
How OpenAI’s disclosure process works
Any OpenAI employee can flag a potential misalignment example and request consideration for public disclosure. The process sets deadlines for each stage.
Technical staff investigate:
- What happened and what remains uncertain
- Whether disclosure is warranted
- Which facts can be shared
- Whether a third party was affected
- Whether private notification is required
Cases follow three tracks:
- Ready for Disclosure: Investigation is sufficiently complete for publication after review.
- Minor Investigation: Further technical investigation is required.
- Larger Investigation (“Slow Track”): A more complex investigation is required, particularly when a third party is involved.
The first two tracks are expected to cover most cases, and all six reports published today fall into them.
Larger investigations and disclosure
For Larger Investigations, security, legal, and responsible disclosure obligations take precedence when a third party is affected. OpenAI aims to provide an initial notice as soon as possible but may delay disclosure for security reasons, such as when a model discovers a previously unknown vulnerability in widely used software.
If a report identifies a third party, OpenAI intends to provide advance notice even when no security boundary was crossed. The initial notice will provide a high-level account, whether outside experts are assisting, and an available estimate for the final report.
OpenAI said the OpenAI Hugging Face incident would have fallen under this track if it had been disclosed under the new framework.
Unresolved disagreements over disclosure are referred to OpenAI’s Safety Advisory Group (SAG), which assesses frontier model capabilities and safeguards, oversees the Preparedness Framework, and advises leadership. Disagreements within SAG or staff objections can be escalated to OpenAI leadership.
What each report will include
Reports will cover:
- Observed behavior, severity, and external impact
- Setting and date or date range
- When the behavior was discovered
- Model or models involved
- Resulting harm, where available
- Discovery process and investigation scope
- Implications for alignment research and technical AI safety
- Unanswered questions
- Measures taken or planned
Some measures may not be available when a report is published if the investigation or fix is still underway. For customer deployments, OpenAI will share as much information as customer privacy and contractual obligations allow.
OpenAI said the six reports are an initial set and do not represent the full range or severity of known misalignment cases or ongoing investigations. It plans to continue disclosing cases that meet its criteria, including complex cases requiring longer investigations or third-party coordination.
