OpenAI says Astra is the first model it rates a critical cyber risk

The company says the safeguards are finally good enough to ship the model, and it has not said when shipping day is.

Abstract EMRGNG cover image for a story about OpenAI

OpenAI said on 1 September that Astra, a model it has not yet released, is the first large language model to cross the critical cybersecurity threshold in its Preparedness Framework. Astra scored full marks on ExploitBench, a test of exploiting known vulnerabilities, and in a modified internal test it found and exploited two previously unknown flaws without a person directing each step.

The framework commits OpenAI to holding back a model at that level until it judges the safeguards sufficient. The company said it delayed Astra’s launch over recent weeks to strengthen and test protections against cyber misuse and against the model acting outside its sandbox, and that Astra did not try to escape its test environment in a scenario built to mirror an earlier agent-escape incident. Advanced cyber features will go first to a small group of vetted testers, then to approved defensive users through a programme it calls Daybreak Blue.

This is the first time a major lab has said one of its own systems reached the capability tier its safety policy treats as most dangerous. OpenAI’s argument is that the model which can find flaws at scale can also help defenders patch them first, so a gated release beats no release. Rival labs have described similar capabilities without invoking a formal threshold, which makes Astra a test of whether the framework changes what actually ships.

What is missing is a date. OpenAI said Astra will arrive soon and did not define the word, or say how large the first tester group is or how approved for defensive use gets decided. It has also not published the modified test that produced the two unknown flaws, so the headline result cannot be checked from outside the company.

Read more here.

More from EMRGNG