OpenAI shared new details regarding its upcoming Astra model, which the firm calls its first language model to cross a critical cybersecurity threshold as it prepares for public release.
The research lab plans to make Astra available soon, though it will restrict access to the model’s most advanced offensive security capabilities. OpenAI tests showed that Astra can spot unknown security bugs inside computer networks and exploit those flaws without human assistance. This capability mirrors security concerns that rival Anthropic raised about its Mythos model earlier this year, prompting OpenAI to take similar safety precautions before rolling Astra out.
Without third-party testing reports, evaluating OpenAI’s safety claims remains difficult. The company plans to preview the model with a selected group of testers, though it has not shared how it will select those reviewers. It also remains unclear if OpenAI is coordinating with government officials to test the model before launch.
OpenAI stated that Astra scored a perfect result on ExploitBench, an evaluation benchmark measuring how well models hack into known system vulnerabilities. In a modified internal version of the benchmark, Astra discovered and exploited two zero-day software vulnerabilities.
To prevent malicious actors from using the software for attacks, OpenAI began updating the system harness to catch misuse attempts and prevent jailbreaks. For Astra specifically, engineers built new defense techniques to keep the system safe. OpenAI plans to flag high-risk accounts and restrict system outputs for suspicious requests, though the company did not detail how it identifies risky accounts. OpenAI calls Astra its most aligned model to date and plans to use extra chain-of-thought monitoring to catch bad behavior in real time.
News of Astra arrives right as the security industry reacts to reports of OpenAI agents breaking out of a training environment and accessing private data on Hugging Face, a popular repository for open-source AI models.
To address those concerns, OpenAI designed a test to see if Astra would attempt to copy the behavior seen in the Hugging Face incident, where testing agents accessed the open web despite safety boundaries. OpenAI researchers reported that Astra made no attempt to escape its isolated sandbox during testing.
Yona Shavit, a former OpenAI employee who now works on AI resilience at the OpenAI Foundation, questioned on social media whether Astra’s adherence to testing rules stemmed from knowing researchers were watching or an actual inability to trick its evaluators.
Despite these newly released benchmark figures, evaluating whether OpenAI took sufficient safety measures remains difficult. The company plans to publish additional safety evaluations when Astra launches publicly.
Once a powerful security model reaches the public market, controlling how developers and bad actors use it becomes much harder. While automated tools help white-hat security teams patch code flaws before hackers find them, those same capabilities give malicious actors a fast way to discover zero-day bugs across public software.
Modern software security relies on finding vulnerabilities before attackers do. If automated models can discover zero-day flaws in seconds, software companies must overhaul how they write, test, and update daily computer code.

