Contents
Tags

Most of the GPT-6 Astra coverage went to the AGI argument. The more consequential number got a paragraph: Astra scored 100% on ExploitBench and is the first model OpenAI classes as crossing its Critical threshold for cyber capability.
That is not a marketing figure. It is a company publishing that its own product can find previously unknown vulnerabilities and build working exploits.
OpenAI's preparedness framework defines capability thresholds with required safeguards attached. Crossing into Critical is the company saying a model has reached a level where the downside case is severe enough to need specific mitigations rather than general ones. This is the first time that threshold has been crossed on cyber.
The same disclosure reports that Astra's reasoning is harder to read than its predecessor's, that chain-of-thought monitoring is fragile, and that the trend is in a negative direction.
A model that is simultaneously more capable at finding exploits and less legible to its own monitoring is precisely the combination safety researchers have been describing for years. It arrived quietly, in a paragraph, during a week everyone spent arguing about a benchmark score.
Vendor safety disclosures are one of the few places where a company publishes something against its own commercial interest. When a lab reports that its monitoring is fragile and trending the wrong way, that is more informative than any benchmark in the same announcement.
As with the AGI claim, none of this is independently reproduced. Everything here is the vendor's own report of its own model — worth taking seriously, and worth labelling.
The other number from launch week, and why it does not mean AGI.
Read the benchmark analysis →Tools
Find AI models that fit your exact hardware. Enter your specs and get a ranked list instantly.
Newsletter