Andon Labs benchmarks show OpenAI's GPT-6 Astra topping Claude Fable 5.1 on vending-machine economics and becoming first model to beat human baseline on all five Drone-Bench autonomous drone-navigation subtasks
Andon Labs benchmarked OpenAI's GPT-6 Astra against Claude Fable 5.1 (and others) on two independent benchmarks: Vending-Bench 2 (running a simulated vending machine business) and Drone-Bench (writing code to autonomously pilot a DJI Tello EDU drone to find and track a person). Astra averaged $15,515 final bank balance vs Fable's $5,422 across six runs each, avoided losses from prepaying shut-down suppliers, refused a price-fixing proposal from GLM-5.3, and won all three Vending-Bench Arena games it played. On Drone-Bench, Astra's best submissions beat the human-AI baseline on all five subtasks (3D reconstruction, localization, navigation, person detection, tracking) for the first time, though average per-task success rates are inconsistent and full end-to-end reliability remains low (~2.8%).
Entities: OpenAI, GPT-6 Astra, Andon Labs, Claude Fable 5.1, Anthropic, GLM-5.3
0 primary
What happened
Independent lab Andon Labs ran OpenAI's GPT-6 Astra and Anthropic's Claude Fable 5.1 through two benchmarks: Vending-Bench 2 (running a simulated vending machine business) and Drone-Bench (writing code to pilot a DJI Tello EDU drone to find and track a person). Across six runs, Astra averaged a $15,515 final bank balance versus Fable's $5,422, and won all three head-to-head Vending-Bench Arena games it played. On Drone-Bench, Astra's best submissions beat the human-AI baseline on all five subtasks for the first time, including 3D reconstruction, though average success rates were inconsistent and full end-to-end task completion succeeded only about 2.8% of the time.
Why it matters
This gives developers and enterprises a concrete, reproducible data point for comparing models on long-horizon agentic tasks (managing a business, sequencing multi-step physical-world actions), which is useful if you are choosing a model for that kind of work. It is weak evidence of anything beyond that: these are proxy tasks on a toy drone and a simulated shop, not deployed systems, and Astra's "refusal" of a price-fixing proposal rests on just three arena games, too small a sample to call an alignment result. Regulators and lawmakers may seize on the drone framing, but the 2.8% end-to-end reliability figure undercuts any claim of near-term real-world capability.
What is noise
The article's framing, "pilots a surveillance drone and runs a business on its own", overstates what happened: this is a lab benchmark with a $100 educational drone, not a deployed surveillance system or an actual business. Andon Labs has a commercial interest in making its benchmarks newsworthy, and the lawmaker-warning angle is speculative advocacy rather than a finding. The alignment claim (refusing price-fixing) is based on n=3 games and should not be generalised.
Watch next
- 01Whether Andon Labs or others replicate the Drone-Bench and Vending-Bench 2 results with different models or larger sample sizes
- 02Any improvement in Astra's end-to-end Drone-Bench success rate beyond the current 2.8%, which is the real bottleneck for practical use
- 03Whether Anthropic or OpenAI respond with their own benchmark data or dispute the methodology
- 04Any regulatory statements or hearings citing this benchmark specifically, which would signal the lawmaker-attention claim has legs
Coverage
1 storyMore capability signals
Full feed →- Deepseek releases V4.1-Flash, an open-source model that sharply cuts KV cache memory and input-processing compute for AI agents10 Sept 202682
- AI systems outperform expert humans in persuasive communication22 Jun 202681
- Anthropic threat report: Claude abused for malware, drone/missile software, mass surveillance, and industrial-scale distillation by Chinese AI labs11 Sept 202680
- WIRED investigation: Flock Safety's AI person-search tools let police run broad description-based surveillance, with weak guardrails against misuse3 Sept 202680