Signum
Feed
Useful signal5 Sept 2026medium confidence

Artificial Analysis revises its Intelligence Index methodology (v4.2), boosting GPT-6 Astra's relative score after other benchmarks showed it far ahead

Artificial Analysis released version 4.2 of its Intelligence Index benchmark suite: added two new benchmarks (AA-Briefcase, GDP.pdf), dropped GPQA-Diamond (saturated), increased private test data weighting to 40%, fixed scoring errors, and adjusted grading methodology. Under the revised scoring, GPT-6 Astra's score rose to four points above its predecessor (Sol), whereas the prior version had scored it roughly on par. Claude Fable 5.1 remains ranked first, Astra second, Meta third.

CapabilityAccess

Entities: Artificial Analysis, OpenAI, GPT-6 Astra, GPT-6 Sol, Anthropic, Claude Fable 5.1

65Useful signal
1 source
0 primary
Was this useful?
01

What happened

Artificial Analysis released version 4.2 of its widely-cited Intelligence Index benchmark suite: it added two new tests (AA-Briefcase and GDP.pdf), dropped the now-saturated GPQA-Diamond, raised private test data weighting to 40%, and fixed unspecified scoring errors. Under the revised methodology, GPT-6 Astra's score jumped to four points above its predecessor GPT-6 Sol, whereas the previous version had scored the two roughly level. The overall ranking order is unchanged: Claude Fable 5.1 first, Astra second, Meta third.

02

Why it matters

This matters mainly for people who use Artificial Analysis's index to make purchasing or integration decisions, since a re-ranking of a reference benchmark can shift perceived competitive standing without any underlying model, price or product actually changing. Developers, enterprises and investors who cite this index in vendor comparisons should note the goalposts moved, not the product. The real-world impact is limited today because nothing shipped or changed in price or availability, but it does affect the narrative around whether GPT-6 Astra underperformed or was previously undervalued by this specific benchmark.

03

What is noise

The Decoder's framing that the overhaul was a response to "skepticism" about GPT-6 Astra's scoring is the reporter's inference, not Artificial Analysis's stated rationale, which cites benchmark saturation and pace of model releases. There is no primary link to the methodology paper or raw scoring changes in the evidence provided, so the four-point swing cannot be independently verified from this reporting alone, and it is only secondhand relay of Artificial Analysis's own release notes.

04

Watch next

  1. 01Whether Artificial Analysis publishes full v4.2 methodology documentation and raw per-benchmark scores so the four-point Astra shift can be independently checked
  2. 02Whether competing benchmarks (Epoch AI, ARC-AGI-3) or OpenAI's own claims converge with or diverge from the new Astra ranking over the next month
  3. 03Whether Artificial Analysis's already-announced v5 supersedes this interim revision quickly, which would suggest v4.2 was a stopgap rather than a durable methodology fix

Coverage

1 story

More capability signals

Full feed →