Signum
Feed
Useful signal2 Sept 2026medium confidence

OpenAI rates upcoming Astra model as first to reach 'critical' cyber-risk tier, while chain-of-thought safety monitoring it relies on is reportedly becoming less reliable

OpenAI classified its upcoming Astra model as reaching 'critical' cyber capability under its Preparedness Framework (a first), citing benchmark results (full marks on ExploitBench, discovery/chaining of two real zero-day V8 vulnerabilities, sandbox escapes and privilege-escalation exploits in expert-led tests using expanded 'Daybreak Blue' access) and released safety figures (91.5% refusal rate on disallowed cyber requests vs 59% for GPT-5.6 Sol; no infrastructure-compromise attempts in a honeypot test vs 56% for predecessor). Advanced cyber features will roll out first to alpha testers, then more broadly via Daybreak Blue for defensive use. Separately, reporting (via The Information) revealed Astra uses a 'recurrent depth' latent-reasoning technique that partially moves reasoning into non-readable internal representations, raising concerns about the reliability of chain-of-thought safety monitoring; OpenAI's chief scientist publicly acknowledged CoT monitoring is 'fragile.'

CapabilityGovernanceInfrastructure

Entities: OpenAI, Astra, GPT-5.6 Sol, Anthropic, Claude Fable 5.1, Mythos 5.1

62Useful signal
1 source
0 primary
Was this useful?
01

What happened

OpenAI has classified its unreleased Astra model as reaching the "critical" tier for cyber capability under its own Preparedness Framework, a first for the company. It cites internal benchmark results (full marks on ExploitBench, chaining of two real V8 zero-days, sandbox escapes in expert-led tests) alongside safety figures showing a 91.5% refusal rate on disallowed cyber requests versus 59% for GPT-5.6 Sol, and zero infrastructure-compromise attempts in a honeypot test versus 56% for its predecessor. Advanced cyber features will initially go to alpha testers, then to a wider group via "Daybreak Blue" access. Separately, reporting from The Information says Astra uses a "recurrent depth" reasoning technique that moves some reasoning into non-readable internal states, and OpenAI's chief scientist has publicly called chain-of-thought monitoring "fragile."

02

Why it matters

If accurate, this is a genuine milestone: a frontier lab self-certifying a model capable of autonomously finding and exploiting unknown software vulnerabilities, which will shape how enterprises, defenders and regulators think about AI-assisted cyberattacks. The gated rollout (alpha, then Daybreak Blue) suggests OpenAI expects real misuse risk, which matters for security teams and policymakers building threat models now. The more durable story may be the chain-of-thought monitoring problem: if the main tool used to check what these models are "thinking" is becoming less reliable just as capability increases, that is a structural oversight gap relevant to every major lab, not just OpenAI.

03

What is noise

Every figure here comes from OpenAI's own framing or secondhand reporting; there is no primary evidence link, published system card, or independent verification of the benchmark claims. The tests were run under conditions OpenAI itself designed and controlled, so "critical capability" and "safest model" claims should be read as marketing-adjacent self-assessment, not third-party audit. The announcement's timing alongside competitor releases (Anthropic's Claude Fable 5.1, Mythos 5.1) is a packaging tell, and headlines framing Astra as simultaneously "most dangerous" and "safest" are doing rhetorical work to generate attention as much as informing risk assessment.

04

Watch next

  1. 01Publication of Astra's official system card and whether independent researchers (e.g. UK AI Security Institute) can reproduce or validate the ExploitBench and zero-day chaining results
  2. 02Whether Daybreak Blue access controls hold up in practice, or whether jailbreaks/leaks expose the advanced cyber capabilities to a wider pool of bad actors within the first few months
  3. 03Further detail or independent scrutiny of the 'recurrent depth' technique and its effect on chain-of-thought monitoring reliability, including whether other labs disclose similar architectural shifts

Coverage

1 story

More capability signals

Full feed →