“These results, in addition to expert assessments, have led us to conclude last night that we cannot rule out critical cyber capabilities under our Preparedness Framework⁠.”

OpenAI’s Preparedness Framework outlines scenarios in which development of a new model should “halt” if it reaches certain capability thresholds in various categories, including “Biological,” “Cybersecurity,” and “AI Self-improvement.”

For cybersecurity, the “critical” threshold means a model can pinpoint “zero-day exploits of all severity levels” in “hardened real-world systems” without any human help. 

A model could also hit the “critical” level if it can carry out “end-to-end novel strategies for cyberattacks against hardened targets” with little more than a “high-level desired goal” in mind, according to the OpenAI safety framework.

OpenAI’s previous high-end model, GPT-5.6 Sol, only reached the “high” threshold during internal evaluations, the company said. OpenAI initially released GPT-5.6 Sol to just a “select group of trusted partners” before making the model public a couple of weeks later.

Given its concerns over Astra’s potential cybersecurity risks, OpenAI says is it “implementing stricter security controls” for the model, such as setting up “isolated testing environments” and “restricted network and tool access,” among other measures. 

Share.
Leave A Reply

Exit mobile version