OpenAI releases GPT-6 Astra system card: fewer hallucinations, near-perfect direct prompt injection defense, but indirect prompt injection failure rate still 8.5%
OpenAI published the GPT-6 Astra system card detailing safety and security benchmark results: reduced hallucination rates vs GPT-5.6 Sol, 99.99% defense against direct prompt injection, 91.5-98.3% jailbreak refusal on single-turn attacks (dropping to ~67% defense over multi-turn adaptive attacks), and an 8.5% failure rate on indirect/hidden prompt injections (down from 27% for GPT-5.6 Sol), per external testing by Gray Swan using 1,810 curated attacks from its IPI Arena.
Entities: OpenAI, GPT-6 Astra, GPT-5.6 Sol, Gray Swan, Claude Opus 5, Anthropic
0 primary
What happened
OpenAI published the system card for GPT-6 Astra, its newest model, alongside external red-team results from Gray Swan (1,810 attacks via its IPI Arena benchmark). Reported figures: hallucination rates are down from GPT-5.6 Sol, direct prompt injection defense is at 99.99%, single-turn jailbreak refusal sits at 91.5-98.3%, and indirect (hidden) prompt injection failures fell to 8.5% from 27% on the prior model. Multi-turn adaptive jailbreak attacks still succeed roughly a third of the time, cutting defense to about 67%.
Why it matters
This gives enterprise security teams a concrete, comparable data point when deciding whether to deploy GPT-6 Astra in agents that read external documents, browse the web, or call tools autonomously. An 8.5% indirect injection failure rate is a real improvement but still means roughly 1 in 12 hidden attacks gets through, which is not a safe bar for high-autonomy, high-volume agent deployments. The multi-turn jailbreak weakness (33% success rate for attackers) is arguably the more important number for anyone building conversational agents that operate over extended sessions.
What is noise
The headline framing of "fewer hallucinations, near-perfect injection defense" glosses over the fact that the strong numbers (99.99% direct defense) apply to the easy case, while the hard case relevant to real deployments (indirect injection, multi-turn jailbreaks) still fails at meaningful rates. Comparisons to Anthropic's Claude Opus 5 (4.8% IPI failure) are based on testing done in a different quarter than Gray Swan's Astra evaluation, so the head-to-head gap is less solid than it looks. These are also bare-model results without the production safety classifiers OpenAI would layer on in a real deployment, so real-world failure rates could be lower, or the comparison could be misleading in the other direction if competitors' numbers include their classifiers.
Watch next
- 01Whether OpenAI or third parties publish post-launch indirect injection incident rates once GPT-6 Astra is deployed in production agents, not just lab benchmarks
- 02Whether Gray Swan or another red team re-tests Claude Opus 5 and GPT-6 Astra in the same quarter under identical conditions to settle the comparison
- 03Whether the multi-turn jailbreak defense rate (currently ~67%) improves in a future point release or system card update, since this is the more operationally relevant weakness for long-running agents
Coverage
1 storyMore capability signals
Full feed →- AI systems outperform expert humans in persuasive communication22 Jun 202681
- WIRED investigation: Flock Safety's AI person-search tools let police run broad description-based surveillance, with weak guardrails against misuse3 Sept 202680
- Hcompany open-sources NeoMME, a from-scratch multimodal-native encoder family, and NeoMME-Retriever for visual document retrieval3 Sept 202679
- Benchmark results show significant improvement in AI agent performance on WorkBench15 Jun 202679