Signum
Feed
Useful signal6 Sept 2026high confidence

Third-party benchmark: OpenAI's GPT-6 Astra outperforms Claude Fable 5.1 at robotic pick-and-place task, ties on harder puzzle-insertion task

An independent evaluator ran OpenAI's GPT-6 Astra on the same YAM robot arms under the same Inspect Robots agent policy used in a prior Claude Fable 5/5.1 comparison, testing two manipulation tasks (block-into-bowl, puzzle-piece-into-groove) with 20 trials each. Astra achieved 19/20 (95%) completion on the bowl task vs Fable 5.1's 8/20 (40%) and Fable 5's 1/20 (5%), at lower cost ($0.94/run) and less time (2.5 min/run) than Fable 5.1 ($2.12/run, 6.8 min). On the puzzle task, Astra tied Fable 5.1 at 2/20 (10%) completions, both stalling at the same final insertion step, though Astra was cheaper ($1.36 vs $2.18/run).

CapabilityEconomics

Entities: OpenAI, GPT-6 Astra, Claude Fable 5, Claude Fable 5.1, Anthropic, Inspect Robots

75Useful signal
1 source
0 primary
Was this useful?
01

What happened

An independent evaluator ran OpenAI's GPT-6 Astra and Anthropic's Claude Fable 5.1 through identical robot-arm manipulation tests using the same YAM hardware and Inspect Robots agent framework, 20 trials per task. On a simple block-into-bowl task, Astra completed 19/20 (95%) versus Fable 5.1's 8/20 (40%) and Fable 5's 1/20 (5%), and did so cheaper ($0.94 vs $2.12 per run) and faster (2.5 vs 6.8 minutes). On a harder puzzle-piece insertion task, both Astra and Fable 5.1 tied at 2/20 (10%), failing at the same final step, though Astra was again cheaper.

02

Why it matters

This gives developers and enterprises evaluating agentic robot control a rare head-to-head with real numbers rather than vendor claims, useful for model selection on simple grasp-and-place tasks where Astra's advantage looks real and substantial. The shared failure on precision insertion is arguably the more important finding: it suggests a genuine capability ceiling in current models for fine manipulation, not just a Claude-specific weakness, which matters for anyone planning near-term deployment of these models in physical automation.

03

What is noise

The "2.4x better" framing is true for exactly one narrow task on one robot rig with 20 trials per condition, a small sample that does not establish general manipulation superiority. This is also an incremental follow-up to an earlier Fable-only report, not a fresh independent study design, and lab pick-and-place results on a single arm setup may not generalise to other tasks, hardware, or real-world conditions.

04

Watch next

  1. 01Whether independent replications on different robot hardware or task sets confirm Astra's bowl-task advantage or narrow it
  2. 02Whether OpenAI or Anthropic respond with their own benchmark data or dispute the methodology
  3. 03Whether either model shows improvement on the puzzle-insertion task in future point releases, which would indicate the capability ceiling is moving

Coverage

1 story

More capability signals

Full feed →