← Blog/Have We Reached AGI With GPT-6 Astra? The 37-Point Gap Says No
deep-dive
Runyard Team
@runyard_dev
8 min read

Tags

#gpt-6-astra#agi#openai#benchmarks#arc-agi#analysis
Runyard.dev — Find AI Models That Run on Your Hardware

Have We Reached AGI With GPT-6 Astra? The 37-Point Gap Says No

OpenAI's 99.9% ARC-AGI score against ARC Prize's 62.7% on a neutral harness
Same model, same benchmark, 37 points apart. The difference is the harness.

OpenAI launched GPT-6 Astra on 3 September and Greg Brockman called it the start of AGI. VentureBeat ran “Welcome to the AGI era”. If you only read headlines, something historic happened this week.

There is a specific, checkable number underneath the claim, and it does not say what the headlines say. That number is the entire post.

The 37-point gap

OpenAI reports GPT-6 Astra scoring 99.9% on ARC-AGI — the benchmark built specifically to measure the kind of general reasoning that people mean when they say AGI. That figure was produced using a new provider adapter harness, which is OpenAI's own scaffolding around the model.

ARC Prize, which runs the benchmark, scored the same model at 62.7% on its provider-neutral harness. State of the art, comfortably — and 37 points below the number in the announcement.

Both figures are real. Neither party is lying. They are measuring different things: one measures a model wrapped in bespoke scaffolding, the other measures the model on the same footing as everything else. The 37 points in between were contributed by the harness.

This is the same argument we made about coding agents, with a much bigger number attached: the harness is often worth more than the model. When the scaffolding can move a score by 37 points, the score describes the scaffolding as much as the intelligence.

Why the tooling around a model often matters more than which model it is.

Read the harness argument

So has anything actually generalised?

That is the question the gap makes hard to answer, and it is the honest answer to “have we reached AGI”. Adding elaborate scaffolding makes it harder, not easier, to tell whether the underlying model has generalised — because you can no longer separate what the model worked out from what the harness handed it.

A system that scores 99.9% with bespoke tooling and 62.7% without has demonstrated something impressive about the system. It has not demonstrated that the model reasons generally. Those are different claims, and the announcement collapsed them.

It is also worth noting that by at least one reading, the AGI claim fails OpenAI's own published bar. The company set thresholds; this release has been reported as not clearing them.

What did genuinely change

Dismissing the release because the headline oversells it would be its own mistake. Several things here are real and significant:

  • Computer use. Astra can operate a machine directly, which is a step change in what an agent can be pointed at rather than a benchmark artefact.
  • Cyber capability. It is the first model OpenAI classes as crossing its Critical threshold — it can find unknown vulnerabilities and build working exploits, scoring 100% on ExploitBench. That is not a marketing number; it is a risk disclosure.
  • Autonomy. Longer, less supervised task execution than Sol managed.

The uncomfortable half of the same disclosure: Astra's reasoning is harder to read than its predecessor's, chain-of-thought monitoring is described as fragile, and the trend is in the wrong direction. A model that is more capable and less legible at the same time is the combination safety people have been warning about.

What would settle it

The question “is this AGI” is unanswerable as posed, because there is no agreed threshold. But the narrower question — has this model generalised beyond its training in a way previous ones did not — has a clear test:

  • Provider-neutral harnesses, reported as the headline number rather than a footnote. If the model is that good, it does not need bespoke scaffolding to prove it.
  • Scores on benchmarks published after the model's training cutoff, where contamination cannot explain the result.
  • Independent replication. Every figure in circulation today traces back to the vendor or to the benchmark's own team.
  • Sustained performance on long, messy, real tasks — not a 30-minute evaluation with a scoring rubric.

The answer

No, and the more interesting fact is that the gap between the two scores is more informative than either score on its own. It tells you that in September 2026, the scaffolding around a model can be worth 37 points on the benchmark specifically designed to be scaffolding-proof.

If you build with these models, that is the practical takeaway and it has nothing to do with AGI: the harness you put around a model may matter more than which model you picked. That was true for coding agents last month and it is true at the frontier this month.

One thing this site can say with more confidence than most: none of the above is arithmetic. Everywhere else on Runyard the numbers are computed and you can check them. Benchmark scores are reported numbers from interested parties, and they should be read that way — including when they support a conclusion you like.

Meanwhile: what the models you can actually run at home need in memory.

Check your own hardware

RUNYARD.DEV

Hardware-aware AI model discovery. Know exactly what runs on your machine — before you download.

© 2026 RUNYARD.DEV — All rights reserved.

Built for local AI.

Tools

Try Runyard

Find AI models that fit your exact hardware. Enter your specs and get a ranked list instantly.

Newsletter