Main page » GPT-6 Astra: What Is New in OpenAI’s Next-Gen Model

On September 3, 2026, OpenAI released GPT-6 Astra, its new flagship model and the successor to GPT-5.6 Sol. The company calls it “the world’s most intelligent and aligned model,” and president Greg Brockman went further at the launch briefing, suggesting it is “not unreasonable to feel that we are now in the AGI era.” The benchmarks are genuinely striking, but so are the caveats behind the headline numbers. Here is what is actually new, what the scores mean, and how to access Astra.

Article preview graphic titled "GPT-6 ASTRA: OPENAI'S NEXT-GENERATION AI MODEL", featuring benchmark score metrics, functional icons, and futuristic digital globe graphics over a dark blue background.

What Is New in GPT-6 Astra

Astra is built for sustained, multi-step work rather than single answers. Its focus is agentic tasks: software development, research, web browsing, computer control, and document creation, where the model keeps working over long horizons instead of replying once and stopping.

The technical specs back the ambition. The GPT-6-Astra API model supports a 1.05-million-token context window, up to 128,000 output tokens, and reasoning levels from low through max. It reportedly came out of OpenAI’s largest training run ever, over 100,000 GPUs at the Stargate site in Texas, and is notably the first OpenAI release where earlier models helped supervise the training of the new one.

The Benchmark Headlines

Astra posts large gains across reasoning, math, coding, and computer use:

  • ARC-AGI-3: a headline 98.6% in OpenAI’s own testing, up from 7.8% for GPT-5.6 Sol six months earlier.

  • FrontierMath Tier 4: 97.6%, a very high score on advanced mathematics.

  • ExploitBench: 100%, a difficult cybersecurity challenge, versus 78.5% for Sol.

  • OSWorld 2.0 (computer use): 72.6%, up from 65.7% for Sol, and completed in roughly 40 minutes on average versus about 75 for its predecessor.

  • Terminal-Bench Science: 64.6%, a large jump from Sol’s 22.4%.

The computer-use gains matter most in practice: as models spend more time operating software and navigating interfaces, speed and reliability count as much as raw reasoning, and Astra improved on both.

The Caveats Behind the Numbers

Here is where honest coverage differs from a press release. The headline ARC-AGI-3 score comes with a real asterisk, and independent testers flagged it.

OpenAI ran Astra through its Responses API harness with settings that preserve reasoning between turns, while the comparison models were often evaluated under different setups. Because ARC-AGI-3 tests how a model navigates an unfamiliar interactive environment, the harness it runs in can significantly change the result. The independent ARC Prize organization tested Astra with its own standardized, provider-neutral setup and recorded 62.7%, still state-of-the-art, but far below the 98.6% to 99.9% OpenAI reported with its own tooling. In other words, both numbers are real; they measure different things. OpenAI’s figure reflects Astra plus its agent system, not the raw model alone.

The coding claims deserve similar care. OpenAI describes Astra as its best software-engineering model yet, but on the public DeepSWE leaderboard it sits in a close cluster with Gemini and Claude Opus 5 rather than clearly leading, and rival models top some coding benchmarks outright. Astra is a major step forward, but it does not dominate every category the way a headline reading might suggest.

Is This “AGI”?

Brockman’s AGI framing made news, but he stopped short of claiming Astra is AGI outright, calling the concept itself “a gray, fuzzy thing.” The ARC Prize team, whose benchmark is designed to measure the gap to AGI, explicitly cautioned against reading a high score as proof the gap has closed. Notably, OpenAI CEO Sam Altman also signaled that future releases will be paced by safety considerations rather than raw capability. The honest summary: Astra is a remarkable model and a real step change in interactive reasoning, but “the AGI era” is a marketing frame, not a settled fact.

Availability and Access

GPT-6 Astra is rolling out now. It reached a limited set of organizations first and, over the following days, became available to all ChatGPT Plus, Pro, Business, and Enterprise users, plus developers through the OpenAI API, Microsoft Azure, and AWS Bedrock. As a premium flagship, it comes at a higher price than previous models, so for many everyday tasks the cheaper-tier models may still be the practical choice.

The Takeaway

GPT-6 Astra is a genuine leap in agentic AI: stronger reasoning, dramatically better computer use, and a scale of training OpenAI has not attempted before. It is also a case study in reading benchmarks carefully, since the most eye-catching number depends heavily on the testing setup, and independent results tell a more measured story. Astra is very likely the most capable model available as of its launch. Whether it marks the arrival of AGI is a much bigger claim than the scores alone can support.

❓ Frequently Asked Questions

Answers to relevant questions about this AI tool

What is GPT-6 Astra?
GPT-6 Astra is OpenAI’s flagship AI model released on September 3, 2026, succeeding GPT-5.6 Sol. It is built for agentic, multi-step work like coding, research, and computer use, and OpenAI describes it as its most intelligent and aligned model to date.
What did GPT-6 Astra score on ARC-AGI-3?
OpenAI reported 98.6% on ARC-AGI-3 using its own Responses API harness, and up to 99.9% at high reasoning, a dramatic jump from GPT-5.6 Sol’s 7.8%. However, the independent ARC Prize measured 62.7% with a standardized provider-neutral setup, so the score depends heavily on the testing configuration.
Is GPT-6 Astra AGI?
Not according to a strict definition. OpenAI president Greg Brockman suggested it may feel like the start of the AGI era, but he did not claim Astra is AGI, and the ARC Prize team cautioned against treating a high benchmark score as proof. It is a major advance, but “AGI” remains a contested, unsettled label.
What are the key specs of GPT-6 Astra?
The gpt-6-astra API model has a 1.05-million-token context window, up to 128,000 output tokens, and reasoning levels from low through max. It was trained on OpenAI’s largest run yet, reportedly over 100,000 GPUs, and targets long-horizon agentic tasks rather than single answers.
How do I access GPT-6 Astra?
Astra is rolling out to ChatGPT Plus, Pro, Business, and Enterprise users, and to developers via the OpenAI API, Microsoft Azure, and AWS Bedrock. It launched to a limited set of organizations first, with wider availability over the following days, at a premium price point.
Is GPT-6 Astra better than Claude or Gemini?
It leads on several reasoning and computer-use benchmarks, but not all. On coding leaderboards like DeepSWE it sits in a close cluster with Claude Opus 5 and Gemini rather than clearly ahead, and rival models top some benchmarks. Astra is a frontier leader, but not a clean sweep across every task.

 

Read more
The week of August 24 to 30 turned on money and law rather than model...
2 weeks ago
0 102
Chrome now has a built-in AI assistant, Gemini! We'll tell you about the new features...
4 weeks ago
0 440
Google Updates Gemini 3.5 with Live Translate Features The release of Gemini 3.5 Live Translate...
4 weeks ago
0 196

Leave a Reply

Your email address will not be published. Required fields are marked *