Signum
Feed
Useful signal24 Sept 2026medium confidence

Epoch AI and MIT studies find AI inference costs for fixed benchmark performance fell sharply since 2023, though estimates of the rate diverge sharply (13x/year vs 3x/year algorithmic-only)

Two research analyses were published/reported comparing AI inference cost trends: Epoch AI finds the price to reach a fixed score on select benchmarks (5 benchmarks spanning math, science, logic) has fallen ~47% per quarter (~13x/year) since 2023, citing OpenAI's o3 matching its early-2025 75% GPQA Diamond score for 1/725th the original cost (30 cents to $0.0004 per question) 18 months later via a GPT-5.6 family model. MIT researchers (Hans Gundlach et al.), using Artificial Analysis pricing data from April 2024-November 2025 across more models, find raw cost declines of 5x-10x/year, but after stripping out cheaper hardware and competitive pricing pressure, estimate pure algorithmic efficiency gains at only ~3x/year. Both also find peak per-run costs can rise because newer reasoning models use more test-time compute, and that benchmark score gains partly reflect more compute spent rather than pure efficiency ("benchmaxxing" concern).

EconomicsCapabilityInfrastructure

Entities: Epoch AI, MIT, Hans Gundlach, OpenAI, o3, GPT-5.6

67Useful signal
1 source
0 primary
Was this useful?
01

What happened

Two separate analyses of AI inference cost trends surfaced via The Decoder, with no primary source links given. Epoch AI reports the price to hit a fixed score on five benchmarks (math, science, logic) has fallen roughly 47% per quarter, or about 13x per year, since 2023, citing OpenAI's o3 matching its early-2025 GPQA Diamond score for 1/725th the cost 18 months later. Separately, MIT's Hans Gundlach and colleagues, using Artificial Analysis pricing data from April 2024 to November 2025, find raw prices falling 5-10x per year, but only about 3x per year once cheaper hardware and competitive pricing are stripped out to isolate pure algorithmic efficiency gains. Both note peak per-run costs can rise because newer reasoning models use more test-time compute.

02

Why it matters

If either figure holds, matching last year's AI capability gets dramatically cheaper over time, which matters for anyone budgeting inference costs, deciding when to upgrade models, or building products that depend on falling unit economics. The 4x gap between Epoch's 13x/year and MIT's 3x/year algorithmic-only estimate matters more than either headline number: it tells you how much of the "efficiency gain" is really just cheaper chips and price wars between vendors rather than genuine algorithmic progress, which affects how durable this trend is. For enterprises and developers, the practical takeaway is that running the current best model can still be flat or rising in cost even as older-tier performance gets cheaper, so budget planning needs to separate "frontier cost" from "commodity cost."

03

What is noise

The "faster than any previous technology" framing and the $50,000-car analogy is marketing packaging, not analysis; cost-curve claims for compute have been made before and Epoch has incentives to produce striking headline multiples. The piece is reported secondhand by The Decoder with no links to the actual Epoch or MIT papers, so the methodology, benchmark selection, and confidence intervals cannot be checked here. Both studies flag a real caveat that gets buried under the framing: benchmark score gains partly reflect spending more compute per query ("benchmaxxing"), which means some of the apparent efficiency gain may just be shifted cost rather than genuine improvement.

04

Watch next

  1. 01Locate and read the actual Epoch AI and MIT (Gundlach et al.) papers to check benchmark selection, sample size, and whether the 13x vs 3x gap is explained or reconciled
  2. 02Track Artificial Analysis pricing data over the next two to three quarters to see which estimate (13x, 5-10x raw, or 3x algorithmic) the trend actually follows
  3. 03Watch whether frontier-model (not fixed-benchmark) inference costs per query rise or fall over the next six months, since that is the number that affects real enterprise budgets rather than backward-looking capability-matching costs

Coverage

1 story

More economics signals

Full feed →