← Blog/Astra Scored 100% on ExploitBench. That Is a Risk Disclosure
deep-dive
Runyard Team
@runyard_dev
7 min read

Tags

#gpt-6-astra#security#openai#safety#exploitbench
Runyard.dev — Find AI Models That Run on Your Hardware

Astra Scored 100% on ExploitBench. That Is a Risk Disclosure

Astra's ExploitBench score and Critical classification
More capable and less legible at the same time.

Most of the GPT-6 Astra coverage went to the AGI argument. The more consequential number got a paragraph: Astra scored 100% on ExploitBench and is the first model OpenAI classes as crossing its Critical threshold for cyber capability.

That is not a marketing figure. It is a company publishing that its own product can find previously unknown vulnerabilities and build working exploits.

What Critical actually means

OpenAI's preparedness framework defines capability thresholds with required safeguards attached. Crossing into Critical is the company saying a model has reached a level where the downside case is severe enough to need specific mitigations rather than general ones. This is the first time that threshold has been crossed on cyber.

The part that should worry you more

The same disclosure reports that Astra's reasoning is harder to read than its predecessor's, that chain-of-thought monitoring is fragile, and that the trend is in a negative direction.

A model that is simultaneously more capable at finding exploits and less legible to its own monitoring is precisely the combination safety researchers have been describing for years. It arrived quietly, in a paragraph, during a week everyone spent arguing about a benchmark score.

What it means if you are not a security researcher

  • Defensive tooling built on frontier models gets better at the same rate as offensive tooling, and both sides now have access to the same capability.
  • Vulnerability disclosure timelines assume finding bugs is slow and expensive. That assumption is weakening.
  • If you run anything internet-facing, the practical change is boring and immediate: patch faster, because the window between a vulnerability existing and being found is shrinking.
  • For local models specifically, nothing changes. Open-weight models at the sizes you can run at home are nowhere near this capability level.

Why this is worth reading carefully

Vendor safety disclosures are one of the few places where a company publishes something against its own commercial interest. When a lab reports that its monitoring is fragile and trending the wrong way, that is more informative than any benchmark in the same announcement.

As with the AGI claim, none of this is independently reproduced. Everything here is the vendor's own report of its own model — worth taking seriously, and worth labelling.

The other number from launch week, and why it does not mean AGI.

Read the benchmark analysis

RUNYARD.DEV

Hardware-aware AI model discovery. Know exactly what runs on your machine — before you download.

© 2026 RUNYARD.DEV — All rights reserved.

Built for local AI.

Tools

Try Runyard

Find AI models that fit your exact hardware. Enter your specs and get a ranked list instantly.

Newsletter