Contents
Tags

OpenAI launched GPT-6 Astra on 3 September and Greg Brockman called it the start of AGI. VentureBeat ran “Welcome to the AGI era”. If you only read headlines, something historic happened this week.
There is a specific, checkable number underneath the claim, and it does not say what the headlines say. That number is the entire post.
OpenAI reports GPT-6 Astra scoring 99.9% on ARC-AGI — the benchmark built specifically to measure the kind of general reasoning that people mean when they say AGI. That figure was produced using a new provider adapter harness, which is OpenAI's own scaffolding around the model.
ARC Prize, which runs the benchmark, scored the same model at 62.7% on its provider-neutral harness. State of the art, comfortably — and 37 points below the number in the announcement.
Both figures are real. Neither party is lying. They are measuring different things: one measures a model wrapped in bespoke scaffolding, the other measures the model on the same footing as everything else. The 37 points in between were contributed by the harness.
This is the same argument we made about coding agents, with a much bigger number attached: the harness is often worth more than the model. When the scaffolding can move a score by 37 points, the score describes the scaffolding as much as the intelligence.
Why the tooling around a model often matters more than which model it is.
Read the harness argument →That is the question the gap makes hard to answer, and it is the honest answer to “have we reached AGI”. Adding elaborate scaffolding makes it harder, not easier, to tell whether the underlying model has generalised — because you can no longer separate what the model worked out from what the harness handed it.
A system that scores 99.9% with bespoke tooling and 62.7% without has demonstrated something impressive about the system. It has not demonstrated that the model reasons generally. Those are different claims, and the announcement collapsed them.
It is also worth noting that by at least one reading, the AGI claim fails OpenAI's own published bar. The company set thresholds; this release has been reported as not clearing them.
Dismissing the release because the headline oversells it would be its own mistake. Several things here are real and significant:
The uncomfortable half of the same disclosure: Astra's reasoning is harder to read than its predecessor's, chain-of-thought monitoring is described as fragile, and the trend is in the wrong direction. A model that is more capable and less legible at the same time is the combination safety people have been warning about.
The question “is this AGI” is unanswerable as posed, because there is no agreed threshold. But the narrower question — has this model generalised beyond its training in a way previous ones did not — has a clear test:
No, and the more interesting fact is that the gap between the two scores is more informative than either score on its own. It tells you that in September 2026, the scaffolding around a model can be worth 37 points on the benchmark specifically designed to be scaffolding-proof.
If you build with these models, that is the practical takeaway and it has nothing to do with AGI: the harness you put around a model may matter more than which model you picked. That was true for coding agents last month and it is true at the frontier this month.
One thing this site can say with more confidence than most: none of the above is arithmetic. Everywhere else on Runyard the numbers are computed and you can check them. Benchmark scores are reported numbers from interested parties, and they should be read that way — including when they support a conclusion you like.
Meanwhile: what the models you can actually run at home need in memory.
Check your own hardware →Tools
Find AI models that fit your exact hardware. Enter your specs and get a ranked list instantly.
Newsletter