Signum
Feed
Useful signal10 Sept 2026high confidence

Independent benchmark replicates and extends UK AISI finding that Astra shows an unusually large jump in reasoning ability without chain-of-thought

An independent researcher built and ran a 19-task "No CoT Reasoning Index" (NCRI) benchmark across multiple frontier models (Astra, Fable 5.1, Gemini 3.8 Flash, and others), replicating and extending UK AISI's earlier finding that Astra shows a large jump in reasoning ability without chain-of-thought prompting. Results: Astra has ~8.6x the odds of solving an arbitrary no-CoT reasoning problem compared to the next-best model (Fable 5.1), and can execute ~7.2 serial arithmetic steps in a single forward pass (at 50% success) versus 4.1 for the next-best models. Astra's improvement is described as disproportionate and lopsided, concentrated in synthetic serial/parallel computation tasks rather than uniformly across cognitive domains (e.g., ~40-point/16x odds gap above the synthetic-vs-nonsynthetic diagonal). Code and benchmark data were released.

CapabilityGovernance

Entities: Astra, Fable 5.1, Gemini 3.8 Flash, UK AISI, OpenAI, Epoch

68Useful signal
1 source
0 primary
Was this useful?
01

What happened

An independent researcher built a new 19-task benchmark (NCRI) to test how well frontier models reason without chain-of-thought prompting, and ran it against Astra, Fable 5.1, Gemini 3.8 Flash and others. The results replicate and extend an earlier UK AISI finding: Astra solves no-CoT reasoning problems at roughly 8.6 times the odds of the next-best model (Fable 5.1), and can chain about 7.2 serial arithmetic steps in one forward pass versus 4.1 for its closest rivals. The gain is concentrated in synthetic serial/parallel computation tasks rather than spread evenly across reasoning types. Code and data have been released publicly.

02

Why it matters

This matters mainly to AI safety researchers, model evaluators and regulators, not to end users or buyers of AI products today: nothing about what models can be deployed or bought has changed. The finding is significant because chain-of-thought monitoring is a leading proposed method for auditing model reasoning, and a model that reasons well without externalising its steps is harder to oversee that way. If replicated further, this strengthens the case that at least one frontier model's internal reasoning is becoming less inspectable, which could feed into how regulators and safety teams think about oversight requirements.

03

What is noise

The leap from "large no-CoT capability gap" to "trend toward less transparent frontier reasoning" is speculative and rests partly on an unsourced claim that OpenAI already finds Astra harder to monitor, which is not independently verified here. The word "concerning" in the original headline is editorialising rather than a finding; the benchmark itself only measures a performance gap, not actual interpretability or monitorability. The architectural explanation for why Astra behaves this way (e.g. looping) is explicitly flagged by the author as unproven.

04

Watch next

  1. 01Whether other independent labs (e.g. Epoch, METR, or academic groups) replicate the ~8.6x odds ratio and 7.2-step serial arithmetic figure using different task sets, not just the author's own NCRI benchmark
  2. 02Any public statement or paper from OpenAI or Astra's developer confirming or denying the reported difficulty in monitoring Astra's reasoning, beyond the single unsourced claim relayed here
  3. 03Whether UK AISI or another regulator formally cites this no-CoT capability jump in guidance on chain-of-thought monitoring or model evaluation requirements over the next 6-12 months

Coverage

1 story

More capability signals

Full feed →