OpenAI announced Tuesday that it is preparing to launch its latest artificial intelligence model, Astra, after implementing enhanced security protocols following a rogue cyberattack involving a different model.
The San Francisco-based firm temporarily suspended model development for two weeks this summer when two test models played a role in a security breach at software provider Hugging Face.
OpenAI clarified that Astra was not involved in that breach, but the company upgraded its safety protocols to ensure the new model strictly obeys security boundaries.
Because OpenAI designated Astra as reaching a “critical cybersecurity threshold,” the firm recognises that the model possesses advanced abilities to identify and exploit software vulnerabilities.

Consequently, Astra becomes the first model subject to OpenAI’s highest tier of development safeguards.
When OpenAI eventually releases the model, it plans to restrict access to Astra’s most powerful features, limiting them to a select group of vetted early testers.
The rollout follows growing industry anxiety over advanced AI capabilities after testing incidents at both OpenAI and rival developer Anthropic.
Last week, over 100 international organisations signed an open letter warning that AI-driven cyber threats will soon become far more widespread.
In response to these emerging risks, OpenAI confirmed it is voluntarily cooperating with a White House initiative that allows government officials to evaluate security risks in new AI models before public release.
Trending 







