OpenAI's GPT-6 Astra hallucinates less but remains vulnerable to hidden prompt injections

Summary
OpenAI's GPT-6 Astra hallucinates less than its predecessor and blocks 99.99 percent of direct prompt injections. But when attacks are hidden inside documents the AI reads, the model still gets cracked in 8.5 percent of scenarios.
Original Article
Captured source content or English translation, normalized into this reading format.
OpenAI's GPT-6 Astra hallucinates less but remains vulnerable to hidden prompt injections
OpenAI's GPT-6 Astra hallucinates less but remains vulnerable to hidden prompt injections

- OpenAI's new model GPT-6 Astra hallucinates less than its predecessor GPT-5.6 Sol and blocks 99.99 percent of direct prompt injection attacks.
- Astra also resists jailbreaks more effectively, but persistent attackers can still get a problematic response about one in three tries over multiple conversation rounds.
- For indirect prompt injections hidden inside documents the AI reads, Astra's failure rate dropped from 27 percent to 8.5 percent. That's a big improvement, but it's still high.
OpenAI's new model, GPT-6 Astra, produces fewer hallucinations and blocks prompt injection attacks more effectively than its predecessors. But it still isn't reliable enough for truly secure AI agent deployments.
Thenew Astra modelmakes far fewer factual errors than its predecessor, GPT-5.6 Sol, according toOpenAI's system card. OpenAI tested it against ChatGPT conversations that users had flagged for wrong answers, meaning these were particularly error-prone cases whose failure rates shouldn't be taken as typical for everyday use. Astra reproduced these reported errors much less often, with the biggest gains showing up at low latency settings and lower reasoning levels.

GPT-6 Astra (pink) hallucinates less than the GPT-5 models across all latency settings. | Image: OpenAI
For direct prompt injections, where users try to manipulate the model through their own prompts, Astra hits a near-perfect 99.99 percent defense rate. OpenAI credits itsGPT-Red methodfor this, which uses an automated attacker to harden the model during training. Ad
Jailbreak resistance looks similar. Against a fixed dataset of known attacks trying to extract harmful responses about biology, violence, and cybersecurity, Astra refuses to help in 91.5 to 98.3 percent of cases. Ad
When attackers adapt their strategy over multiple conversation rounds, Astra's defense rate drops to about 67 percent, meaning persistent adversaries can coax out at least one problematic response roughly one in three tries. Predecessor models scored just under 50 percent on the same test. OpenAI notes that these tests ran on the bare model without the production safety layers like classifiers that ship with the actual product.
Hidden prompt injections remain a real security problem
Astra makes progress on indirect prompt injections, where an attack is buried inside a document the AI reads. External testing by security firm Gray Swan, using 1,810 curated attacks from theirIPI Arena, found that with 15 attempts per scenario, Astra was cracked at least once 8.5 percent of the time. GPT-5.6 Sol failed 27 percent of the time. Claude Opus 5 did better at 4.8 percent in the same evaluation, but it wasn't immune either. Ad

Attack success rates for indirect prompt injections according to Gray Swan's IPI Arena. | Image: OpenAI
The numbers in Gray Swan's combined Q1 and Q2 test actually went up compared to earlier results.Anthropic previously reported only a two percent attack success ratebased on the easier Q1 test alone, and GPT-5.6 Sol scored just 20 percent there too. Anthropic also ran all models with extended reasoning turned on, which, alongside the broader test scope, could explain the gap.
Even though these are curated, hand-picked attacks, the success rates should worry any enterprise security team. Astra can be tricked through injected instructions in roughly one out of every twelve scenarios. Opus 5 holds up better, but it still fails about one in twenty-one. And the risk is growing. AI agents are increasingly writing code, operating tools, and controlling computers on their own, which is what Gray Swan tested. These agents are also being built to run around the clock and at scale, with reading and processing documents as one of their core jobs. Ad
AI News Without the Hype – Curated by Humans
Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.
Region
Europe
Heat Score
81
Category
Global Briefs
Language
en
