GPT-6 Astra ARC-AGI-3: What a Benchmark Number Does Not Establish
Contributed content: this article was written by a third-party contributor and does not necessarily reflect the views of ABC Money. Editorial and Advertising Policy
The GPT-6 Astra figure recorded on ARC-AGI-3 by the independent coverage board, read 2026-10-07, is 62.7%. The same board records OpenAI’s GPT-6 Astra at 95% on ARC-AGI-2 — a gap of thirty-two points between two benchmarks that share a family name and almost nothing else. The endpoint is metered here at list rate with no markup, so nothing below rests on a figure we chose.
That gap is the article. Not because 62.7% is a bad number, but because the two figures are routinely stacked in the same sentence as though they measured the same thing, and the arithmetic that comes out of doing so is not a statement about the model.
Two benchmarks share a name and not a scale
The confusion starts with the naming, which is a habit of the field rather than a fault of any one publication. Successive releases in a benchmark family keep the family name and take a new number, and readers carry the old number’s associations across with them.
The board that records these results keeps them in separate columns with separate leaderboards, and that separation is the honest signal: ARC-AGI-2 and ARC-AGI-3 are different task sets, scored on different scales, populated by different runs. A score is a property of *a model on a version*, never of the model alone. The version is not a suffix.
Two consequences follow, and both catch people who work quickly.
A cross-version comparison is not a comparison. “Model A scores 95 and Model B scores 63” invites the reading that A is far stronger. Where the two numbers come from different versions, that reading is unsupported — you are comparing two different tests that happen to be filed under adjacent headings.
A version change is indistinguishable from a capability change on a page that does not print the version. This is the ordinary failure mode of a benchmark table that travels: the version is context, and context is the first thing dropped when a number is quoted.
The thirty-two-point cliff
Put the four models in this comparison side by side on both versions, and the shape of the record matters more than any single cell:
| (benchlm.ai, read 2026-10-07) | ARC-AGI-2 | ARC-AGI-3 |
| GPT-6 Astra | 95% | 62.7% |
| GPT-6.1 Sol | 94.2% | 52.7% |
| Claude Opus 5.5 | 91.7% | no row |
| Grok 4.7 | no row | no row |
Read it by row and then by column. By row, each model that appears twice drops sharply between versions — Astra by 32 points, Sol by 41. By column, the ordering on the older benchmark is nearly a tie across three models, while on the newer one the spread widens to ten points between the two OpenAI rungs.
Neither reading says anything about general reasoning. What the table does say is that the number moves with the version, and it moves by more than the differences between the models. Any argument built on the distance between two models on this benchmark is being made inside the noise of a version change, and most published arguments of that shape make the comparison across versions without noticing.
There is a second reason to be careful, and it is structural rather than arithmetic. A single pass rate is one number on one task family. It cannot show where a model’s reasoning breaks down, how it degrades as tasks get longer, or whether it succeeds for the intended reason rather than by a shortcut the task designer did not anticipate. A benchmark with a headline pass rate has one dimension; a claim about reasoning has several.

What a saturated score would establish
Set the recorded figure aside for a moment and take the harder question head-on, because it is the one readers actually arrive with: if a model posts a score at or close to saturation on a version of ARC-AGI-3, does that mean it is generally intelligent?
No, and the reason is not that the benchmark is easy. It is that a pass rate is silent about the thing the question is about.
It does not establish generality. A benchmark is a finite task set, and a very high pass rate on it says the model handles that set well. Whether the set is a fair sample of the cognitive territory is the benchmark author’s claim, not a measurement, and it is the claim a saturated score makes most tempting to skip past.
It does not establish that the success is the intended one. High scores on puzzle-shaped benchmarks have repeatedly turned out to include solutions found by routes the designers did not anticipate. A pass rate counts outcomes; it does not audit the process that produced them, and a benchmark that does not publish its failing cases makes that audit impossible from the outside.
It does not establish that the number is the model’s. Every score is taken at a configuration. On the independent evaluation harness, Artificial Analysis measures this generation of models at a named reasoning effort — Max for GPT-6 Astra and GPT-6.1 Sol, Xhigh for Grok 4.7, read 2026-10-07. A model at its top reasoning setting is being given more computation per task than the same model at a lower rung, and a table that omits the setting is comparing configurations, not models.
And it does not survive the version boundary. This is the one that matters most in practice. A near-saturated score on one version, quoted next to a mid-sixties score on the next, produces the appearance of a collapse or a triumph depending on which way the sentence points. Both readings are wrong; the version changed.
The second number nobody puts beside it
There is a second published measurement for this model that belongs in the same paragraph as its reasoning score, and it is almost never quoted next to one.
The same independent board publishes a prompt-injection robustness figure — an attack success rate under Gray Swan’s injection evaluation, where lower is better, read 2026-10-07. GPT-6 Astra records 8.5%. Claude Opus 5.5 records 1.0%. GPT-6.1 Sol and Grok 4.7 carry no row.
Read that honestly in both directions. On this axis Astra is roughly eight times more susceptible than the model that tops the same set, and a system that pipes untrusted text into a model with tool access is exposed to that difference in a way that a reasoning score never shows. What it does *not* mean is that 8.5% is a threshold — the figure is one evaluation’s attack set, and the metric is binary per attempt, so the gap is a relative signal between two models rather than a probability about your application.
The useful habit is to keep the two axes together. A model can lead a reasoning table and sit mid-pack on robustness, because the two are measured by different harnesses against different objectives, and neither one is a summary of the other. A stack that only tracks the reasoning column is tracking half of what the independent boards publish about the models it depends on.

How to read an ARC-AGI-3 score in the wild
Four checks, in the order that catches the most errors.
Find the version before the number. ARC-AGI-2 and ARC-AGI-3 are separate tests with separate scales. A page that prints a score without the version has told you less than it appears to, and a page that prints both without labelling them has told you something misleading.
Find the effort setting before comparing two models. Astra and Sol are measured at Max, Grok 4.7 at Xhigh. A cross-model table without those labels is a table of configurations.
Ask what the score does not cover. A pass rate says nothing about injection resistance, nothing about long-context degradation, and nothing about the failure modes your own workload will meet. On this model, the robustness row is published on the same page as the reasoning row and is quoted far less often.
Make the last check your own. Every figure above is a third-party reading, and third-party readings are re-run. The only degradation curve that describes your traffic is the one you plot on your own documents — a hundred real questions at several context depths will tell you more than a leaderboard row, and it costs an afternoon.
The takeaway
GPT-6 Astra records 62.7% on ARC-AGI-3 and 95% on ARC-AGI-2, both on benchlm.ai and both read 2026-10-07 — a thirty-two-point gap that is a property of the benchmark version rather than of the model. A near-saturated score on one version of a benchmark is not evidence of general intelligence, and it is not comparable to a mid-sixties score on the next version; the task set, the scale and the leaderboard all changed with the number.
The same board publishes a prompt-injection figure that belongs beside the reasoning figure and rarely appears there: 8.5% for Astra against 1.0% for Claude Opus 5.5, lower being better. Two axes, two harnesses, and no summary of one in the other.
Read every ARC-AGI-3 score as three facts stacked in one number — the model, the version, and the configuration — and quote all three or none.
OrcaRouter carries GPT-6 Astra at list rate on the same key as the rest of the OpenAI line, so a version-controlled comparison of your own workload against either rung is a model-string change rather than a second integration.
Sourcing note: the ARC-AGI-2 and ARC-AGI-3 figures (95% and 62.7% for GPT-6 Astra, 94.2% and 52.7% for GPT-6.1 Sol, 91.7% for Claude Opus 5.5 on ARC-AGI-2), the Gray Swan prompt-injection rates (8.5% for GPT-6 Astra, 1.0% for Claude Opus 5.5) and the absence of rows for GPT-6.1 Sol and Grok 4.7 are from benchlm.ai’s live model pages, read 2026-10-07. The Max reasoning setting recorded for GPT-6 Astra and GPT-6.1 Sol and the Xhigh setting recorded for Grok 4.7 are from Artificial Analysis’ live model pages, read the same day. The reading of the ARC-AGI-2-to-ARC-AGI-3 gap as a property of the benchmark version rather than of the model, the observation that the version change exceeds the between-model spread, and the list of what a pass rate does not establish are this article’s own analysis of those published figures, not published claims. The 8.5%-versus-1.0% comparison is stated as a relative signal between two models on one evaluation’s attack set and is not a probability estimate for any application. No vendor claim is cited in this article. All checked 2026-10-07; re-check before 2026-11-01.