GPT-6 Astra Release: 99.9% ARC-AGI-3 Score
OpenAI shipped GPT-6 Astra today — the frontier model yesterday's Path to Astra post said crossed a "critical cybersecurity capability threshold." It rolls out to a limited set of organizations now, then all ChatGPT Plus, Pro, Business, and Enterprise users, the API, and AWS over the coming days.
What shipped
Astra saturates ARC-AGI-3 at 99.9% — against 7.8% for GPT-5.6 Sol. ARC Prize independently verified it: 99.9% at high reasoning effort for ~$19K of inference, with Astra beating the median human on action efficiency in 96% of levels and building its own compact symbolic world-models on the fly. It also hits 100% on ExploitBench (Sol: 78.5%), 97.6% on FrontierMath Tier 4, and 72.6% on OSWorld 2.0 at ~40 minutes per task — 47% faster than Sol. API pricing: gpt-6-astra at $10/M input, $50/M output, with a fast mode at 2.5x speed for 2x price.
Why it matters
The ARC-AGI-3 number is the story. Oracle covered ARC-AGI-1 falling to 44% in 67 cents a month ago; the residual gap keeps collapsing. But this is the model OpenAI delayed after the Hugging Face incident — and the release notes confirm why: 100% ExploitBench, two previously unknown zero-days found during evaluation, SRE-Bench at 88% vs Sol's 55.9%. It ships with refusal training for PoC exploits and misalignment monitoring, with less restricted safeguards coming via Daybreak. The alignment numbers hold the line — 0% out-of-scope behavior vs Sol's 48% — but the contain-the-frontier question just moved into production.