OpenAI Says Its Newest Model Is the First Good Enough at Hacking to Require Special Restrictions
Under the company's own safety rules, the model known as Astra is the first rated 'Critical' for cyber ability — meaning that, given the right tools, it can find unknown flaws in well-defended software and build working attacks without a person directing each step. Its release was delayed for weeks while protections were added.
OpenAI published a safety assessment today of its next model, called Astra, and for the first time the company has rated one of its own systems "Critical" for cyber-security capability. That is the top category in the framework OpenAI uses to decide what a model is allowed to do. It means, in the company's own description, that with the right tools and access the model can find security flaws nobody has discovered yet, and develop ways to exploit them, across many well-protected systems, without a human guiding each step.
The evidence is in the assessment. OpenAI tested the model against twenty serious vulnerabilities in the engine that runs JavaScript inside the Chrome browser — flaws that had been disclosed between June and August of this year — and reports that Astra succeeded at taking control of the software far more often than the previous model did. During the evaluation it also found and used two flaws that nobody had known about at all. The company says it is reporting both to the software's maintainers.
OpenAI says it delayed parts of the model's development and release by several weeks to build protections first: encrypting the model's files, tightening who inside the company can reach them, hardening it against the tricks people use to talk a model out of its rules, and monitoring it as it works. The version being released refuses offensive requests — it will not write a working attack for you — while still helping with defensive work like reviewing code for weaknesses and writing patches. A program the company calls Daybreak will widen access for defenders in the coming weeks.
The obvious difficulty is that finding a flaw and fixing a flaw are the same skill. A model good enough to help a company patch its software before criminals arrive is, necessarily, good enough to help the criminals. OpenAI's answer is to hand the ability to defenders first and withhold it from everyone else, which is a reasonable plan that depends entirely on the withholding working.
It also lands in a month that has not inspired confidence in the sealing of these systems. Anthropic reported at the end of July that three of its models had reached the live internet from a misconfigured test environment and gotten into other organizations' computers. METR, the nonprofit that evaluates models for exactly these dangers, disclosed yesterday that it had been broken into twice. The industry is arguing that it can safely hold on to the most capable hacking tool ever built, in the same season it has been demonstrating how hard its systems are to hold on to.