Signum
Feed
Useful signal13 Sept 2026high confidence

Andon Labs benchmarks show OpenAI's GPT-6 Astra topping Claude Fable 5.1 on vending-machine economics and becoming first model to beat human baseline on all five Drone-Bench autonomous drone-navigation subtasks

Andon Labs benchmarked OpenAI's GPT-6 Astra against Claude Fable 5.1 (and others) on two independent benchmarks: Vending-Bench 2 (running a simulated vending machine business) and Drone-Bench (writing code to autonomously pilot a DJI Tello EDU drone to find and track a person). Astra averaged $15,515 final bank balance vs Fable's $5,422 across six runs each, avoided losses from prepaying shut-down suppliers, refused a price-fixing proposal from GLM-5.3, and won all three Vending-Bench Arena games it played. On Drone-Bench, Astra's best submissions beat the human-AI baseline on all five subtasks (3D reconstruction, localization, navigation, person detection, tracking) for the first time, though average per-task success rates are inconsistent and full end-to-end reliability remains low (~2.8%).

CapabilityEconomicsPowerGovernance

Entities: OpenAI, GPT-6 Astra, Andon Labs, Claude Fable 5.1, Anthropic, GLM-5.3

62Useful signal
1 source
0 primary
Was this useful?
01

What happened

Independent lab Andon Labs ran OpenAI's GPT-6 Astra and Anthropic's Claude Fable 5.1 through two benchmarks: Vending-Bench 2 (running a simulated vending machine business) and Drone-Bench (writing code to pilot a DJI Tello EDU drone to find and track a person). Across six runs, Astra averaged a $15,515 final bank balance versus Fable's $5,422, and won all three head-to-head Vending-Bench Arena games it played. On Drone-Bench, Astra's best submissions beat the human-AI baseline on all five subtasks for the first time, including 3D reconstruction, though average success rates were inconsistent and full end-to-end task completion succeeded only about 2.8% of the time.

02

Why it matters

This gives developers and enterprises a concrete, reproducible data point for comparing models on long-horizon agentic tasks (managing a business, sequencing multi-step physical-world actions), which is useful if you are choosing a model for that kind of work. It is weak evidence of anything beyond that: these are proxy tasks on a toy drone and a simulated shop, not deployed systems, and Astra's "refusal" of a price-fixing proposal rests on just three arena games, too small a sample to call an alignment result. Regulators and lawmakers may seize on the drone framing, but the 2.8% end-to-end reliability figure undercuts any claim of near-term real-world capability.

03

What is noise

The article's framing, "pilots a surveillance drone and runs a business on its own", overstates what happened: this is a lab benchmark with a $100 educational drone, not a deployed surveillance system or an actual business. Andon Labs has a commercial interest in making its benchmarks newsworthy, and the lawmaker-warning angle is speculative advocacy rather than a finding. The alignment claim (refusing price-fixing) is based on n=3 games and should not be generalised.

04

Watch next

  1. 01Whether Andon Labs or others replicate the Drone-Bench and Vending-Bench 2 results with different models or larger sample sizes
  2. 02Any improvement in Astra's end-to-end Drone-Bench success rate beyond the current 2.8%, which is the real bottleneck for practical use
  3. 03Whether Anthropic or OpenAI respond with their own benchmark data or dispute the methodology
  4. 04Any regulatory statements or hearings citing this benchmark specifically, which would signal the lawmaker-attention claim has legs

Coverage

1 story

More capability signals

Full feed →