Signum
Feed
Useful signal4 Sept 2026high confidence

OpenAI releases GPT-6 Astra system card: fewer hallucinations, near-perfect direct prompt injection defense, but indirect prompt injection failure rate still 8.5%

OpenAI published the GPT-6 Astra system card detailing safety and security benchmark results: reduced hallucination rates vs GPT-5.6 Sol, 99.99% defense against direct prompt injection, 91.5-98.3% jailbreak refusal on single-turn attacks (dropping to ~67% defense over multi-turn adaptive attacks), and an 8.5% failure rate on indirect/hidden prompt injections (down from 27% for GPT-5.6 Sol), per external testing by Gray Swan using 1,810 curated attacks from its IPI Arena.

CapabilityGovernanceAdoption

Entities: OpenAI, GPT-6 Astra, GPT-5.6 Sol, Gray Swan, Claude Opus 5, Anthropic

77Useful signal
1 source
0 primary
Was this useful?
01

What happened

OpenAI published the system card for GPT-6 Astra, its newest model, alongside external red-team results from Gray Swan (1,810 attacks via its IPI Arena benchmark). Reported figures: hallucination rates are down from GPT-5.6 Sol, direct prompt injection defense is at 99.99%, single-turn jailbreak refusal sits at 91.5-98.3%, and indirect (hidden) prompt injection failures fell to 8.5% from 27% on the prior model. Multi-turn adaptive jailbreak attacks still succeed roughly a third of the time, cutting defense to about 67%.

02

Why it matters

This gives enterprise security teams a concrete, comparable data point when deciding whether to deploy GPT-6 Astra in agents that read external documents, browse the web, or call tools autonomously. An 8.5% indirect injection failure rate is a real improvement but still means roughly 1 in 12 hidden attacks gets through, which is not a safe bar for high-autonomy, high-volume agent deployments. The multi-turn jailbreak weakness (33% success rate for attackers) is arguably the more important number for anyone building conversational agents that operate over extended sessions.

03

What is noise

The headline framing of "fewer hallucinations, near-perfect injection defense" glosses over the fact that the strong numbers (99.99% direct defense) apply to the easy case, while the hard case relevant to real deployments (indirect injection, multi-turn jailbreaks) still fails at meaningful rates. Comparisons to Anthropic's Claude Opus 5 (4.8% IPI failure) are based on testing done in a different quarter than Gray Swan's Astra evaluation, so the head-to-head gap is less solid than it looks. These are also bare-model results without the production safety classifiers OpenAI would layer on in a real deployment, so real-world failure rates could be lower, or the comparison could be misleading in the other direction if competitors' numbers include their classifiers.

04

Watch next

  1. 01Whether OpenAI or third parties publish post-launch indirect injection incident rates once GPT-6 Astra is deployed in production agents, not just lab benchmarks
  2. 02Whether Gray Swan or another red team re-tests Claude Opus 5 and GPT-6 Astra in the same quarter under identical conditions to settle the comparison
  3. 03Whether the multi-turn jailbreak defense rate (currently ~67%) improves in a future point release or system card update, since this is the more operationally relevant weakness for long-running agents

Coverage

1 story

More capability signals

Full feed →