On September 3, 2026, OpenAI released GPT-6 Astra, its new flagship model and the successor to GPT-5.6 Sol. The company calls it “the world’s most intelligent and aligned model,” and president Greg Brockman went further at the launch briefing, suggesting it is “not unreasonable to feel that we are now in the AGI era.” The benchmarks are genuinely striking, but so are the caveats behind the headline numbers. Here is what is actually new, what the scores mean, and how to access Astra.

What Is New in GPT-6 Astra
On this page
Astra is built for sustained, multi-step work rather than single answers. Its focus is agentic tasks: software development, research, web browsing, computer control, and document creation, where the model keeps working over long horizons instead of replying once and stopping.
The technical specs back the ambition. The GPT-6-Astra API model supports a 1.05-million-token context window, up to 128,000 output tokens, and reasoning levels from low through max. It reportedly came out of OpenAI’s largest training run ever, over 100,000 GPUs at the Stargate site in Texas, and is notably the first OpenAI release where earlier models helped supervise the training of the new one.
The Benchmark Headlines
Astra posts large gains across reasoning, math, coding, and computer use:
-
ARC-AGI-3: a headline 98.6% in OpenAI’s own testing, up from 7.8% for GPT-5.6 Sol six months earlier.
-
FrontierMath Tier 4: 97.6%, a very high score on advanced mathematics.
-
ExploitBench: 100%, a difficult cybersecurity challenge, versus 78.5% for Sol.
-
OSWorld 2.0 (computer use): 72.6%, up from 65.7% for Sol, and completed in roughly 40 minutes on average versus about 75 for its predecessor.
-
Terminal-Bench Science: 64.6%, a large jump from Sol’s 22.4%.
The computer-use gains matter most in practice: as models spend more time operating software and navigating interfaces, speed and reliability count as much as raw reasoning, and Astra improved on both.
The Caveats Behind the Numbers
Here is where honest coverage differs from a press release. The headline ARC-AGI-3 score comes with a real asterisk, and independent testers flagged it.
OpenAI ran Astra through its Responses API harness with settings that preserve reasoning between turns, while the comparison models were often evaluated under different setups. Because ARC-AGI-3 tests how a model navigates an unfamiliar interactive environment, the harness it runs in can significantly change the result. The independent ARC Prize organization tested Astra with its own standardized, provider-neutral setup and recorded 62.7%, still state-of-the-art, but far below the 98.6% to 99.9% OpenAI reported with its own tooling. In other words, both numbers are real; they measure different things. OpenAI’s figure reflects Astra plus its agent system, not the raw model alone.
The coding claims deserve similar care. OpenAI describes Astra as its best software-engineering model yet, but on the public DeepSWE leaderboard it sits in a close cluster with Gemini and Claude Opus 5 rather than clearly leading, and rival models top some coding benchmarks outright. Astra is a major step forward, but it does not dominate every category the way a headline reading might suggest.
Is This “AGI”?
Brockman’s AGI framing made news, but he stopped short of claiming Astra is AGI outright, calling the concept itself “a gray, fuzzy thing.” The ARC Prize team, whose benchmark is designed to measure the gap to AGI, explicitly cautioned against reading a high score as proof the gap has closed. Notably, OpenAI CEO Sam Altman also signaled that future releases will be paced by safety considerations rather than raw capability. The honest summary: Astra is a remarkable model and a real step change in interactive reasoning, but “the AGI era” is a marketing frame, not a settled fact.
Availability and Access
GPT-6 Astra is rolling out now. It reached a limited set of organizations first and, over the following days, became available to all ChatGPT Plus, Pro, Business, and Enterprise users, plus developers through the OpenAI API, Microsoft Azure, and AWS Bedrock. As a premium flagship, it comes at a higher price than previous models, so for many everyday tasks the cheaper-tier models may still be the practical choice.
The Takeaway
GPT-6 Astra is a genuine leap in agentic AI: stronger reasoning, dramatically better computer use, and a scale of training OpenAI has not attempted before. It is also a case study in reading benchmarks carefully, since the most eye-catching number depends heavily on the testing setup, and independent results tell a more measured story. Astra is very likely the most capable model available as of its launch. Whether it marks the arrival of AGI is a much bigger claim than the scores alone can support.
❓ Frequently Asked Questions
Answers to relevant questions about this AI tool