Independent benchmark replicates and extends UK AISI finding that Astra shows an unusually large jump in reasoning ability without chain-of-thought
An independent researcher built and ran a 19-task "No CoT Reasoning Index" (NCRI) benchmark across multiple frontier models (Astra, Fable 5.1, Gemini 3.8 Flash, and others), replicating and extending UK AISI's earlier finding that Astra shows a large jump in reasoning ability without chain-of-thought prompting. Results: Astra has ~8.6x the odds of solving an arbitrary no-CoT reasoning problem compared to the next-best model (Fable 5.1), and can execute ~7.2 serial arithmetic steps in a single forward pass (at 50% success) versus 4.1 for the next-best models. Astra's improvement is described as disproportionate and lopsided, concentrated in synthetic serial/parallel computation tasks rather than uniformly across cognitive domains (e.g., ~40-point/16x odds gap above the synthetic-vs-nonsynthetic diagonal). Code and benchmark data were released.
Entities: Astra, Fable 5.1, Gemini 3.8 Flash, UK AISI, OpenAI, Epoch
0 primary
What happened
An independent researcher built a new 19-task benchmark (NCRI) to test how well frontier models reason without chain-of-thought prompting, and ran it against Astra, Fable 5.1, Gemini 3.8 Flash and others. The results replicate and extend an earlier UK AISI finding: Astra solves no-CoT reasoning problems at roughly 8.6 times the odds of the next-best model (Fable 5.1), and can chain about 7.2 serial arithmetic steps in one forward pass versus 4.1 for its closest rivals. The gain is concentrated in synthetic serial/parallel computation tasks rather than spread evenly across reasoning types. Code and data have been released publicly.
Why it matters
This matters mainly to AI safety researchers, model evaluators and regulators, not to end users or buyers of AI products today: nothing about what models can be deployed or bought has changed. The finding is significant because chain-of-thought monitoring is a leading proposed method for auditing model reasoning, and a model that reasons well without externalising its steps is harder to oversee that way. If replicated further, this strengthens the case that at least one frontier model's internal reasoning is becoming less inspectable, which could feed into how regulators and safety teams think about oversight requirements.
What is noise
The leap from "large no-CoT capability gap" to "trend toward less transparent frontier reasoning" is speculative and rests partly on an unsourced claim that OpenAI already finds Astra harder to monitor, which is not independently verified here. The word "concerning" in the original headline is editorialising rather than a finding; the benchmark itself only measures a performance gap, not actual interpretability or monitorability. The architectural explanation for why Astra behaves this way (e.g. looping) is explicitly flagged by the author as unproven.
Watch next
- 01Whether other independent labs (e.g. Epoch, METR, or academic groups) replicate the ~8.6x odds ratio and 7.2-step serial arithmetic figure using different task sets, not just the author's own NCRI benchmark
- 02Any public statement or paper from OpenAI or Astra's developer confirming or denying the reported difficulty in monitoring Astra's reasoning, beyond the single unsourced claim relayed here
- 03Whether UK AISI or another regulator formally cites this no-CoT capability jump in guidance on chain-of-thought monitoring or model evaluation requirements over the next 6-12 months
Coverage
1 storyMore capability signals
Full feed →- AI systems outperform expert humans in persuasive communication22 Jun 202681
- WIRED investigation: Flock Safety's AI person-search tools let police run broad description-based surveillance, with weak guardrails against misuse3 Sept 202680
- Google DeepMind launches AlphaGenome Atlas, a free public database of predicted effects for 9 billion possible human genome variants8 Sept 202679
- Hcompany open-sources NeoMME, a from-scratch multimodal-native encoder family, and NeoMME-Retriever for visual document retrieval3 Sept 202679