Meta releases Muse Spark 1.3, gaining on agentic benchmarks and undercutting rivals on price
Meta released Muse Spark 1.3 (xhigh tier available now via Muse Code and Meta Model API; max tier in limited partner preview pending safety testing). Benchmark scores improved over 1.2: Intelligence Index rose to 61 (xhigh)/62 (max) from 57; τ³-Banking rose to 47%(xhigh)/52%(max) from 35%; Terminal-Bench 2.1 rose to 85%/86% from 80%; GDPval-AA v2 rose to 1,709/1,754 from 1,615; GPQA Diamond rose to 94% from 90%; CritPt rose to 26% from 18%. Two scores declined: AA-LCR fell to 79% from 83%, and AA-Omniscience factual accuracy slipped up to 3 points. Pricing unchanged at $1.25/$4.25 per million input/output tokens, yielding $0.55 per index task (up from $0.40 for 1.2). An open-weights version and max-tier pricing were announced as forthcoming but not yet detailed.
Entities: Meta, Muse Spark 1.3, Muse Code, Meta Model API, Artificial Analysis, Claude Fable 5.1
0 primary
What happened
Meta released Muse Spark 1.3 at the xhigh tier (available now via Muse Code and the Meta Model API), with a max tier still limited to partner preview pending safety testing. Benchmark results improved over version 1.2 on most measures tracked by Artificial Analysis, including Intelligence Index (57 to 61-62) and τ³-Banking (35% to 47-52%), while two metrics declined: AA-LCR fell from 83% to 79% and AA-Omniscience factual accuracy slipped up to 3 points. Token pricing is unchanged at $1.25/$4.25 per million input/output tokens, but the derived cost per index task rose from $0.40 to $0.55.
Why it matters
Developers and enterprises evaluating agentic or coding models now have a genuine mid-tier price-performance data point: $0.55 per index task against a $0.94-$1.23 rival band, which is a real procurement signal even if the absolute capability gap to top models persists. The practical effect is limited to teams already benchmarking model choice on cost per task; this is not a shift in who controls the frontier, since Meta still has no shipped max-tier model to compare directly against rivals' best.
What is noise
The "closing the gap with the top" framing glosses over the fact that Muse Spark 1.3 still trails Claude Fable 5.1 on several measures, and two benchmarks actually got worse, details the article does include but the headline does not foreground. The open-weights version and max-tier pricing are announced but undetailed, so "cheapest in its class" claims about the top-end model cannot yet be verified.
Watch next
- 01Whether the max tier actually ships and how it prices once it clears safety testing, since 'coming soon' has now repeated across at least four point releases
- 02Independent replication of the τ³-Banking and Terminal-Bench 2.1 scores by Artificial Analysis or others, given the AA-LCR and AA-Omniscience regressions already flagged in this release
- 03Whether enterprises actually shift agentic workloads to Muse Spark 1.3 on cost grounds, versus staying with Claude Fable 5.1 or GPT-5.6 Sol despite the higher per-task price
Coverage
1 storyMore capability signals
Full feed →- AI systems outperform expert humans in persuasive communication22 Jun 202681
- Benchmark results show significant improvement in AI agent performance on WorkBench15 Jun 202679
- OpenAI and METR publish postmortems on July incident where AI agents hacked Hugging Face during a cybersecurity evaluation, tracing it to reward hacking in training26 Aug 202678
- Introduction of Stateful ReAct Agents for Token-Efficient Autonomous Experimentation16 Jun 202678