xAI releases Grok 4.7, a cheaper coding/knowledge model that trails Claude and GPT-6 on independent benchmarks
xAI released Grok 4.7, built on a larger base model with longer reinforcement learning training, priced at $2/million input tokens and $6/million output tokens, and made available via the Grok API, Cursor, and Grok Build. Independent benchmarking (Artificial Analysis Intelligence Index v4.3.2) scores it 46 overall (mid-pack) versus 53 for both Claude Fable 5.1 and GPT-6, and on Terminal-Bench 4.0 agentic coding it scores 26%, behind GPT-6 Astra (60%), Claude Fable 5.1 (55%), and even DeepSeek V4.1 Flash (27%).
Entities: xAI, Grok 4.7, Elon Musk, Claude Fable 5.1, GPT-6, GPT-6 Astra
0 primary
What happened
xAI launched Grok 4.7, priced at $2 per million input tokens and $6 per million output tokens, available through the Grok API, Cursor, and Grok Build. On the Artificial Analysis Intelligence Index (v4.3.2), it scores 46 overall against 53 for both Claude Fable 5.1 and GPT-6. On Terminal-Bench 4.0, a coding-agent benchmark, it scores 26%, behind GPT-6 Astra (60%), Claude Fable 5.1 (55%), and even the cheaper DeepSeek V4.1 Flash (27%).
Why it matters
Anyone choosing a model for coding or agentic tasks now has a concrete price-versus-capability data point: Grok 4.7 is cheap but currently the weakest of the named frontier and near-frontier models on independent coding benchmarks. It is not a reason to switch away from Claude or GPT-6 for serious coding work, but it may be relevant for low-stakes, high-volume tasks where cost matters more than accuracy. For xAI, this is evidence the company has not closed the capability gap despite a bigger base model and longer RL training.
What is noise
xAI's framing of Grok 4.7 as its "most capable model yet" and emphasis on "improved self-verification" are marketing claims not tested by the cited benchmarks. The benchmark figures come from Artificial Analysis via The Decoder, with no direct links to the xAI announcement or the benchmark pages in the underlying data, so these numbers cannot be independently re-verified from this article alone. Low pricing is being framed as a competitive strength, but a 20-34 point benchmark deficit against similarly priced or cheaper alternatives undercuts that pitch.
Watch next
- 01Whether Artificial Analysis or another independent benchmarker publishes updated or disputing scores for Grok 4.7 in the coming weeks
- 02Adoption signals: whether Cursor, Grok Build, or other coding tools report meaningful usage share shifting to Grok 4.7 given its price
- 03xAI's next model release timeline and whether it narrows the Terminal-Bench and Intelligence Index gap versus Claude and GPT-6, which would indicate a real capability trajectory rather than a one-off cheap release
Coverage
1 storyMore capability signals
Full feed →- Deepseek releases V4.1-Flash, an open-source model that sharply cuts KV cache memory and input-processing compute for AI agents10 Sept 202682
- Anthropic threat report: Claude abused for malware, drone/missile software, mass surveillance, and industrial-scale distillation by Chinese AI labs11 Sept 202680
- WIRED investigation: Flock Safety's AI person-search tools let police run broad description-based surveillance, with weak guardrails against misuse3 Sept 202680
- Google DeepMind launches AlphaGenome Atlas, a free public database of predicted effects for 9 billion possible human genome variants8 Sept 202679