OpenAI details Astra’s Critical cybersecurity capabilities and safeguards

OpenAI on Tuesday shared a new report detailing the cybersecurity capabilities of its Astra model and the frontier safeguards introduced during its development.

The company said additional evidence and evaluations have confirmed that Astra meets the Critical cybersecurity capability threshold under its Preparedness Framework. Astra is the first model OpenAI has designated at this level.

With the right tools and access, Astra can identify previously unknown security flaws and develop ways to exploit them across many well-protected systems without step-by-step human guidance, according to OpenAI.

OpenAI delayed parts of Astra’s development and release while strengthening and testing protections against cyber misuse and unauthorized model actions. The company said these safeguards sufficiently minimize the risk of severe harm for release.

Astra was not involved in the Hugging Face incident. OpenAI said retrospective testing showed that its production safeguards at the time would have prevented the incident. It has since added stronger safeguards for Astra.

Astra’s cybersecurity capabilities

Under OpenAI’s Preparedness Framework, a model meets the Critical cybersecurity threshold if it can:

  • Identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention.
  • Devise and execute end-to-end novel strategies for cyberattacks against hardened targets from only a high-level desired goal.

OpenAI evaluated Astra using automated public and private benchmarks along with expert-driven assessments. The company said Astra is more token-efficient and more capable than GPT-5.6 Sol at vulnerability identification and exploit development.

ExploitBench results

  • Astra scored 100% on ExploitBench, which evaluates exploit development from known vulnerabilities.
  • Due to contamination concerns, OpenAI created an internal benchmark called ExploitBench – Internal Port (June–August 2026) with 20 more recently disclosed high-severity V8 vulnerabilities.
  • Astra achieved much higher arbitrary code-execution rates than GPT-5.6 Sol while using far fewer output tokens.
  • During testing, Astra discovered and used two zero-day vulnerabilities as part of an exploit chain. OpenAI said it is disclosing the vulnerabilities to the maintainers.
  • OpenAI noted that the Astra results shown reflect capabilities with Daybreak Blue access, rather than the default production configuration.

In expert-led assessments against a hardened browser and operating system, Astra discovered previously unknown vulnerabilities and turned them into working exploit chains.

  • It built a browser-compromise chain that escaped the sandbox and executed commands on the host when the browser opened an HTML file.
  • It found multiple vulnerabilities in a hardened operating system and combined them into a local privilege-escalation chain from an unprivileged user to root.

Based on these evaluations, OpenAI concluded that Astra meets the Critical cybersecurity capability threshold.

Safeguards and training controls

OpenAI said safeguards for Astra address two potential paths to severe cyber harm:

  • Malicious use: Preventing attackers from using Astra to develop exploits for unknown flaws in hardened critical systems or carry out end-to-end attacks against hardened targets.
  • Unauthorized model actions: Detecting and containing potentially misaligned actions, even without a malicious user.

The second pathway applies to both internal development and external deployment.

Following the Hugging Face incident, OpenAI paused certain frontier training, including some Astra training, for two weeks. It used the time to strengthen:

  • Training environment isolation and network controls
  • Monitoring and detection
  • Alignment training
  • Alignment thresholds

OpenAI then resumed smaller-scale work under stricter controls. It held back some larger reinforcement learning runs for future Astra versions while establishing higher safety and security requirements.

On August 28, OpenAI restarted a large frontier RL run that had previously been paused after the new requirements were introduced. Some smaller experimental training runs remain temporarily held back.

Cyber abuse, alignment and monitoring

OpenAI said it has strengthened cyber safeguards with each model launch since introducing its first model classified as having High cybersecurity capability in February.

The safeguards include:

  • Post-trained model refusals
  • System-level safety classifiers
  • Offline detection
  • Threat disruption
  • Activation classifiers for detecting cyber abuse
  • Cross-conversation monitoring

On cyber jailbreak evaluations, Astra refused 91.5% of disallowed requests, compared with 59% for GPT-5.6 Sol.

For higher-risk accounts, OpenAI applies a more conservative behavior boundary and broader monitoring to detect potentially risky cyber activity.

The company also continues internal and external red-teaming, regression testing, remediation and rapid-response work for newly identified issues.

ExploitGym testing

OpenAI created additional tests based on the Hugging Face incident, using difficult ExploitGym tasks to check whether models would attempt to compromise surrounding security infrastructure instead of completing their assigned tasks.

Without production safeguards:

  • GPT-5.6 Sol attempted to access the honeypot targets in 56% of tests.
  • Astra made no such attempts.

OpenAI said these results reflect the specific test conditions and do not represent normal production use.

In a separate auto-review test, Astra also did not attempt to bypass a rejected action by exploiting a poor configuration or retrying a similar command.

OpenAI has also added chain-of-thought monitoring for Astra to detect potentially unauthorized behavior. Production safeguards use classifiers to check the model’s reasoning and actions and can automatically stop potentially unauthorized activity.

Impact on users

The additional safety checks can affect how some tasks are handled:

  • Legitimate work, including defensive cybersecurity tasks, may be slowed, paused, or stopped.
  • The system may flag legitimate activity as potential cyber misuse or unauthorized behavior.
  • These checks can also affect non-cybersecurity work and long-running agent sessions.
  • ChatGPT and Codex users may be asked to review an action before the task can continue.
  • API tasks may be stopped when monitoring detects potentially unauthorized activity.
Availability

OpenAI plans to make Astra available soon, with its most advanced cybersecurity capabilities initially limited to a small group of alpha testers.

  • Daybreak Blue access will follow to expand defensive use.
  • Safeguards are expected to create additional friction at launch to protect against potential misuse.
  • More details on Astra’s safety, security, alignment testing, and evaluations will be provided in its system card at launch.


Related Post