GPT-6 Astra gets conflicting benchmark scores from Epoch AI and Artificial Analysis, but posts human-level move-efficiency on ARC-AGI-3, prompting Chollet to move up his AGI forecast
OpenAI's GPT-6 Astra model was benchmarked by Epoch AI, Artificial Analysis, and ARC Prize (ARC-AGI-1/2/3, FrontierMath Erdős). Results diverge: Epoch AI scores it highest overall (169 ECI points vs 267 models); Artificial Analysis scores it 61 (tied with predecessor Sol, behind Claude Fable 5.1's 66). On ARC-AGI-3 (novel game-world reasoning), Astra scored 62.7% on ARC Prize's internal harness (vs 7.8% for Sol, 30.2% for Claude Opus 5), and separately 99.9% using OpenAI's own harness — a gap ARC Prize flags as not a fair cross-vendor comparison. Astra solved 2 of 68 open FrontierMath Erdős problems with Lean-verified proofs. It is 2.5x more expensive per token than Sol (~75% more per task) but cheaper than Claude Opus 5/Fable 5 on comparable coding tasks due to needing far fewer compute steps. ARC Prize documented Astra inventing its own symbolic shorthand notation to track game states, and matching or beating median human move-efficiency on 96% of ARC-AGI-3 levels it solved.
Entities: OpenAI, GPT-6 Astra, GPT-5.6 Sol, Anthropic, Claude Fable 5.1, Claude Opus 5
0 primary
What happened
OpenAI's GPT-6 Astra was benchmarked by three independent evaluators with conflicting results: Epoch AI ranks it first overall among 267 models, while Artificial Analysis scores it level with its predecessor (Sol) and behind Anthropic's Claude Fable 5.1. On ARC-AGI-3, Astra scored 62.7% on ARC Prize's own harness versus 99.9% on OpenAI's harness, a discrepancy ARC Prize explicitly says is not a fair comparison. It also matched or beat median human move-efficiency on 96% of solved ARC-AGI-3 levels, solved 2 of 68 open FrontierMath Erdős problems, and costs 2.5x more per token than Sol while showing regressions on GDPval-AA v2, SciCode, banking support tasks and long-context handling.
Why it matters
For developers and enterprises choosing models, this is genuinely useful signal: concrete cost, token-efficiency and task-specific performance data that affects procurement and deployment decisions right now. But the contradictory rankings between Epoch AI and Artificial Analysis mean no one evaluator's verdict should be taken as definitive, and the documented regressions suggest Astra is a mixed upgrade rather than a clean win over Sol or Claude Opus 5.
What is noise
The framing around François Chollet moving up his AGI timeline is speculation stacked on a measurement, and the ARC-AGI-3 "human-beating efficiency" headline leans on the 62.7% figure while glossing over the 99.9% harness discrepancy that ARC Prize itself flagged as not comparable. The "human-level" and symbolic reasoning language is doing more narrative work than the underlying numbers support, especially given the benchmark disagreement and quiet regressions.
Watch next
- 01Whether independent labs (not just OpenAI's own harness) replicate the 62.7% vs 99.9% ARC-AGI-3 gap, and whether ARC Prize publishes a normalised cross-vendor comparison
- 02Artificial Analysis and Epoch AI updates or reconciliations of their conflicting rankings over the next 2-4 weeks, and whether other evaluators (e.g. LMSYS, Scale) weigh in
- 03Whether OpenAI or third parties address the documented regressions (GDPval-AA v2, banking support, SciCode, long-context) and whether Astra's 2.5x cost premium changes enterprise adoption or procurement decisions
Evidence
1 linkedCoverage
1 storyMore capability signals
Full feed →- AI systems outperform expert humans in persuasive communication22 Jun 202681
- WIRED investigation: Flock Safety's AI person-search tools let police run broad description-based surveillance, with weak guardrails against misuse3 Sept 202680
- Hcompany open-sources NeoMME, a from-scratch multimodal-native encoder family, and NeoMME-Retriever for visual document retrieval3 Sept 202679
- Benchmark results show significant improvement in AI agent performance on WorkBench15 Jun 202679