Signum
Feed
Useful signal10 Sept 2026high confidence

New public dataset OpenDiscoveryTrace releases 558 AI scientist agent trajectories with detailed process traces, revealing frontier model error-rate differences

Researchers released a public dataset, OpenDiscoveryTrace, containing 558 complete AI scientific agent trajectories with structured 9-field-per-step process traces (thoughts, tool calls, observations, errors, revision triggers, confidence) across 124 scientific tasks in drug discovery, materials science, genomics, and literature analysis, covering seven models (GPT-5.4, Claude Opus 4.6, Gemini 3.1 Pro at 124 trajectories each; four open-weight models at 30 each; plus 60 live-retrieval variants), along with trace schema, agent harness, and five defined benchmark tasks with baseline models, released under CC BY 4.0.

CapabilityInfrastructure

Entities: OpenDiscoveryTrace, GPT-5.4, Claude Opus 4.6, Gemini 3.1 Pro, Qwen2.5-7B, Mistral-7B-v0.3

66Useful signal
1 source
0 primary
Was this useful?
01

What happened

Researchers released OpenDiscoveryTrace, a public dataset of 558 complete AI agent trajectories from scientific tasks (drug discovery, materials science, genomics, literature analysis), each with a structured 9-field process trace covering thoughts, tool calls, observations, errors and confidence. It covers seven models, including GPT-5.4, Claude Opus 4.6 and Gemini 3.1 Pro at 124 trajectories each, plus four smaller open-weight models, and is released under CC BY 4.0 with a trace schema, agent harness and five benchmark tasks. The headline finding, drawn from an unreviewed arXiv preprint, is that Claude Opus 4.6 showed roughly 30 times more errors than GPT-5.4 despite similar 84-89% success rates.

02

Why it matters

This gives researchers and auditors a rare look at how AI agents actually work through scientific tasks step by step, not just whether they get the right answer, which matters for anyone trying to diagnose failure modes or build governance tooling around agentic AI. It is a resource for the research and regulatory community rather than something that changes any product, deployment or market position today. Developers evaluating agent reliability now have a concrete public benchmark to test against, but there is no evidence yet that any lab or buyer will act on it.

03

What is noise

The eye-catching 30x error-rate gap rests on only 363 LLM-judged trajectories from a single agent harness that has not been peer reviewed, so the harness design (not just the underlying models) could easily be driving the gap. The claim that this "supports AI governance research" is broader framing than a dataset release actually delivers; it is an evaluation resource, not a governance mechanism.

04

Watch next

  1. 01Whether independent researchers reproduce the Claude Opus 4.6 vs GPT-5.4 error-rate gap using the released traces and a different harness
  2. 02Whether the paper is accepted at a peer-reviewed venue and whether the error-rate finding survives review
  3. 03Whether any AI lab, auditor or regulator publicly cites or builds on the OpenDiscoveryTrace benchmark in the next 3-6 months

Evidence

1 linked

Coverage

1 story

More capability signals

Full feed →