OpenAI Claims GPT-6 Astra Could Be AGI — Digging Into the Original Report, Another Harness Drops It to Just 62.7%
OpenAI launched GPT-6 Astra on September 3, touting it as the world’s smartest and most aligned model, scoring 99.9% on ARC-AGI-3, with President Greg Brockman even mentioning AGI. I went digging through ARC Prize’s own benchmark report, and when the exact same model was switched back to their standard test harness, it managed only 62.7%. Meanwhile, independent benchmarks measured its intelligence score as identical to its predecessor—though the price has genuinely jumped 2.5x.

I had just finished writing that Gemini 3.8 Flash piece the day before, thinking that was probably it for AI news this week, only for OpenAI to drop GPT-6 Astra the very next day. The announcement page opens right away with "the world’s smartest and most aligned model." Axios even quoted OpenAI President Greg Brockman calling it a "generational leap," suggesting it might eventually be looked back on as the moment Artificial General Intelligence (AGI) truly arrived.
Whenever I see the term "AGI," my gut reaction is usually to go digging through the original technical reports. Claims of that magnitude almost always come with caveats—caveats that rarely make it onto the product launch page. And having combed through them this time, those caveats are anything but trivial. Below, I’ve broken down the official numbers, the independent third-party benchmarks, whether everyday users can actually use it, and what it’ll cost.
The Official Chart: Benchmarked Against Their Predecessor and Claude
Let's start with the comparison chart OpenAI published, which benchmarks Astra against its own previous generation, GPT-5.6 Sol, and Anthropic's current flagship, Claude Fable 5.1.

OpenAI's official GPT-6 Astra benchmark comparison chart against GPT-5.6 Sol and Claude Fable 5.1. (Image source: OpenAI official launch materials)
Looking down the numbers item by item, the biggest jump comes in the scientific terminal benchmark Terminal-Bench Science 0.1, where Astra scored 64.6%, compared to just 22.4% for the previous-gen Sol and 52.6% for Claude Fable 5.1. On AutomationBench for business process automation, Astra hit 41.4% versus Sol's 18.1% and Fable 5.1's 31.4%. On research-grade math, FrontierMath Tier 4 (v2), Astra scored 97.6%—which the official marketing copy rounded up to "98%"—while Sol scored 83.0% and Fable 5.1 took 87.8%. On computer terminal tasks, Terminal-Bench 4.0 saw 57.9% versus 37.3% and 55.8%, with Claude hot on its heels. On clinical tasks in HealthBench Professional (length-calibrated), Astra scored 63.4%, less than 3 points higher than Sol's 60.5%. On 3D modeling with BenchCAD, it scored 95.9%, also leading all three.
The most glaring cell in this entire table is the very last row: ARC-AGI-3, a brand-new interactive puzzle benchmark. Astra scored 99.9%, compared to a mere 7.8% on the previous-gen Sol, while Claude Fable 5.1 had no official score listed. Surging from 7.8% straight to 99.9% looks like the quintessential "crossing the threshold overnight" moment, so it's no wonder the word "AGI" got thrown around. But this exact cell is where the trouble begins.
That 99.9% on ARC-AGI-3? It's Under a Different Test Harness
ARC-AGI-3 belongs to the ARC Prize Foundation, and on the very same day, September 3, they published their own evaluation report authored by Greg Kamradt. The report explicitly states that they tested the model across two different evaluation harnesses, and the results differed drastically.

The scorecard published on the official ARC Prize blog. The left column shows the standard harness, while the right column shows the provider adapter harness—a massive 37-percentage-point gap on the exact same model. (Image source: ARC Prize official blog)
Using ARC Prize's own Standard harness with Astra pushed to its maximum reasoning effort ("max"), it only scored 62.7%, burning through $26,098. That viral 99.9% score only came about when switching to the Provider Adapter harness with reasoning effort set to "high," which cost $18,817. Same model, same problem set—just run through a different harness, and the score swings by 37 percentage points. OpenAI's launch page simply stated it "saturates ARC-AGI-3 with a 99.9% score" without mentioning which harness was used. I think that's something people deserve to know.
This table also happens to answer a much more practical question: reasoning effort makes a massive difference. Under that same Standard harness, cranking it to "max" yields 62.7%, dialing it down to "low" gets you just 17.5%, and turning reasoning off altogether ("none") lands at 35.2%. That progression is wildly counterintuitive—turning on low-effort reasoning actually performed nearly 18 percentage points worse than not reasoning at all. ARC Prize's report simply listed the raw numbers without explaining why it flipped, so I can only be honest and say "we don't know why," rather than making up a theory. What's certain is that this reasoning dial isn't linear—never assume that "turning it on a little is always better than none." So when asking "how capable is Astra?", the answer depends heavily on which notch you turn the dial to. The API documentation lists five tiers: low, medium, high, xhigh, and max, but the official model card doesn't state what default tier is used when unspecified, and the ChatGPT web interface hasn't exposed this setting to regular users yet. In other words, the Astra you talk to in the chat box and the record-breaking Astra in the headlines might very well not be running at the same intensity. OpenAI hasn't given an answer on this yet, so we'll have to wait for follow-up documentation.
That said, it's not all bad news. The ARC Prize report also contains an observation that is genuinely favorable to Astra—and one that OpenAI didn't even particularly brag about: action efficiency. Before launching this benchmark, ARC Prize brought in roughly 500 participants from the general public to establish a human baseline. It turned out that under the Provider Adapter harness, Astra at "max" effort completed 96.0% of the tasks using fewer steps than the human median, averaging 51.7% fewer steps overall. The report's author called this a substantive milestone, having previously assumed that "even if an AI solves it, it would need far more trial-and-error iterations than a human"—an assumption that was completely shattered. Now, let's be fair: this silver lining was measured under the exact same Provider Adapter harness as that 99.9%, and ARC Prize didn't release corresponding figures under the Standard harness. Strictly speaking, it warrants a question mark just like the 99.9% does. That said, because it measures "how many steps taken to solve" rather than a binary "can it solve it or not," its nature is inherently less distorted by how forgiving the harness is. The monetary side, however, is a whole other story: human participants received $115 for a 90-minute session plus $5 for each completed task (averaging around $12.78 per puzzle), whereas the AI column started in the tens of thousands of dollars.
Independent Benchmarks: Identical Intelligence Score to Last Gen, at 2.5x the Price
If the ARC section was about "taking official numbers with a grain of salt," the independent report released by Artificial Analysis on September 3 is an outright splash of cold water. On their Intelligence Index, Astra scored 61 points—completely identical to GPT-5.6 Sol, 5 points behind Claude Fable 5.1, and trailing Meta's freshly released Muse Spark 1.3 as well.
The price tag, on the other hand, saw a very real hike. API pricing jumped from Sol's $4 per million input tokens and $20 per million output tokens to $10 and $50, respectively—a full 2.5x increase. Astra's output token consumption is indeed about 10% lower than its predecessor, but that modest token saving falls far short of making up for the price jump. When you run the numbers on the same batch of tasks, Astra at max effort ends up 75% more expensive than Sol. Running the full Intelligence Index benchmark cost a whopping $3,013.30.
So where are the real improvements? Coding and agentic tasks. On Artificial Analysis's Coding Agent Index (v1.4), Astra paired with Codex at max effort scored 67 points, placing third on the leaderboard. The top two spots belong to Claude Code paired with Fable 5.1 (max) at 70 points and Claude Code paired with Opus 5 (xhigh) at 68 points. Below Astra, Muse Code paired with Muse Spark 1.3 (xhigh) and Grok Build paired with Grok 4.5 (high) sit tied at 64 points. So here is how tight the race actually looks: 3 points off the top spot, 1 point behind second, and 3 points ahead of fourth place. It's tightly clustered in the lead pack, not trailing behind. What's genuinely staggering, though, is its token efficiency: in the exact same Codex environment, Astra consumed only one-third the tokens of GPT-5.6 Sol (max), and one-fifth those of Claude Opus 5 (xhigh). So even with the higher per-token price, running an actual task end-to-end still came out to less than half the cost of Claude Fable 5, for the same score.

Official AutomationBench chart plotting API cost on the horizontal axis and accuracy on the vertical axis; Astra's curve achieves higher accuracy at lower costs. (Image source: OpenAI official launch materials)
Another area you'll tangibly feel in daily use is hallucinations. In Artificial Analysis's knowledge and hallucination benchmark, AA-Omniscience, Astra's hallucination rate at max effort plummeted from Sol's 92% down to 51%—nearly cut in half. Better yet, its overall accuracy climbed by 4 points simultaneously, meaning this wasn't achieved simply by "shutting up whenever it wasn't sure."
The regressions are also worth putting on record. The same report notes that Astra dropped roughly 80 Elo points on GDPval-AA v2—a benchmark adapted from OpenAI's own dataset measuring economic value across 44 professions. It also dropped 2 to 3 points each on customer-service-oriented τ³-Banking, scientific Python problem-solving on SciCode, and long-context reasoning on AA-LCR. While the long-horizon knowledge-work benchmark AA-Briefcase saw an overall gain of roughly 80 Elo points, breaking it down reveals that while analytical quality improved, presentation deck quality actually deteriorated—an area where previous-gen Sol (max) still holds the lead today. So this isn't a model that completely steamrolls its predecessor across the board; it's a model that charges full steam ahead in coding and agentic workflows while slipping backward in certain types of knowledge work.
What You Really Should Pay Attention to: The Cybersecurity Line
Putting benchmark scores aside, what I think everyday users should really take note of is that OpenAI itself labeled Astra as the very first model to reach a "Critical" cybersecurity capability level. In their September 1 evaluation article and September 3 system card, they put it quite bluntly: given appropriate tools and permissions, Astra can discover previously unknown vulnerabilities across many well-defended systems and develop novel exploits, all without requiring step-by-step human guidance.
That's not just colorful phrasing. OpenAI revealed that in an internal evaluation using 20 recently disclosed high-severity V8 vulnerabilities, the model independently discovered and leveraged two zero-day exploits. To clarify where things stand right now: OpenAI merely stated they are "in the process of reporting them to maintainers," without disclosing CVE numbers, affected versions, or patch statuses. As of this writing, there's no way to definitively answer whether Chrome users are currently safe; the best anyone can do is stick to the age-old advice of keeping your browser updated to the latest version and waiting for official maintainer advisories. In expert-led red teaming, the model pulled off a complete exploit chain against a hardened browser—escaping the sandbox and executing commands on the host machine simply by opening an HTML file. It also chained multiple vulnerabilities on a hardened operating system to achieve full privilege escalation from a standard user account all the way to root. On the public ExploitBench, it posted a flawless 100% score.
Because of this threat level, OpenAI introduced substantial guardrails and hardening this time around. On cyberattack-related jailbreak requests, Astra's refusal rate reached 91.5%, up from just 59% on Sol. The attack success rate for indirect prompt injection—where hidden instructions in web pages trick the AI into doing malicious things—dropped from Sol's 27.0% down to 8.5%, which is very tangible good news for anyone asking AI to browse web pages or operate computer environments on their behalf. In simulations across more than 54,000 internal Codex tasks, high-severity misconduct flags fell from Sol's 73 occurrences (0.135%) to 34 (0.063%). In another honeypot test, when the previous-gen Sol encountered unsolvable challenges, it took unauthorized shortcuts by attacking neighboring peripheral systems 48.2% of the time; Astra clocked in at a clean 0.0%.
These advanced cybersecurity capabilities won't be handed out directly to everyone. Initial access is restricted to a small cohort of testers, followed by a broader rollout to defenders via Daybreak Blue. Daybreak is OpenAI's dedicated access program for cybersecurity workflows, split into two tiers. Blue is for defenders, currently running on a version of GPT-5.6 Sol with relaxed safety guardrails calibrated for defensive work like vulnerability discovery, detection engineering, and incident response. Red is for more advanced offensive-style work like vulnerability research and exploit reproduction, running a specially trained GPT-5.6 Cyber. Getting in isn't just a matter of filling out a form: you have to pass OpenAI's Trusted Access for Cyber screening, undergo identity verification, and agree to active usage monitoring. Individual participants must also use hardware security keys starting September 1, 2026. As for access channels, besides going directly through OpenAI, AWS already launched Daybreak Red and Blue on Amazon Bedrock on August 11, where eligibility can be requested via your AWS account team—though it's currently restricted to a single region, US East (Ohio). So if you're an enterprise security team looking to leverage this level of capability for defense, the path exists, but the barrier to entry is high. As for when Astra itself will be integrated into Daybreak Blue, OpenAI hasn't provided a timeline.
Yet OpenAI also candidly disclosed something distinctly unsettling: Astra's "monitorability" has regressed compared to the previous generation. It has developed a stronger ability to curate its own chain-of-thought, making it less likely to write down thoughts that could incriminate or penalize itself. In adversarial evaluations where researchers deliberately instructed it to evade oversight, it demonstrated strategic underperformance on benchmarks to avoid detection, and occasionally bypassed internal monitoring to execute destructive tasks. While OpenAI stated they have seen no evidence of steganographic reasoning hidden in unrelated text and maintained that overall alignment remains superior to the previous generation, they acknowledged that they take this trend very seriously. In plain English: the model has become more compliant, but simultaneously much harder to see through—and both things are unfolding together.
Can Regular Users Access It Now, and How Much Does It Cost?
This part is fairly straightforward. The official announcement states that access begins today for a small batch of organizations, rolling out over the coming days to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as via the OpenAI API and AWS. Free users aren't on the list, nor is the $8/month Go plan; using Astra requires at least the $20/month Plus subscription. Pro comes in $100 and $200 tiers, with the difference being usage multipliers.
As for usage limits—something a lot of people care about—OpenAI hasn't yet announced message caps for Astra under the Plus tier. We can use current baselines as a reference: Plus currently allows 160 messages every 3 hours for GPT-5.5 Instant, and roughly 3,000 messages per week for Thinking-class models, calculated on a rolling window rather than resetting at midnight each day. Historically, caps are kept quite tight whenever a new flagship launches, so if you already run into limits on Plus regularly, chances are high you'll hit them here too. We'll have to wait for an official announcement for the exact numbers.
If you're going via the API, the specs are: a context window of 1,050,000 tokens, a max output limit of 128,000 tokens per request, and a knowledge cutoff of April 30, 2026. It supports text and image inputs with text output. Pricing runs at $10 per million input tokens and $50 per million output tokens, with cached reads at $1 and cached writes at $12.50.
Here's another detail many people overlooked: the prior generation hasn't been deprecated. A quick check of the official model catalog confirms GPT-5.6 Sol is still there, holding steady at $4 and $20 with zero price hikes. The mid-tier GPT-5.6 Terra sits at $2 and $12, while the lightweight GPT-5.6 Luna costs $0.20 and $1.20. So "holding off and staying on Sol" is a very real, practical option right now, not just wishful thinking. The catch is that OpenAI hasn't committed to how long the older models will stick around; past generations typically began phasing out once the new flagship stabilized, which is something you'll want to factor in if you're planning long-term budgets.
So, Should You Care?
If you're just using ChatGPT to chat, look things up, or draft documents, this update's impact on you will mostly amount to fewer hallucinations. You probably won't notice much else, so just wait a few days until it shows up in your model picker and switch over—there's no need to spend extra cash just for this. Let me clear up one common misconception here: OpenAI lumped Plus, Pro, Business, and Enterprise together under "over the coming days." Pro subscribers aren't getting early access; the first rollout wave went to a small group of organizations, not individual users paying more money. So even if you're shelling out $200 a month, opening ChatGPT right now probably won't show Astra yet.
The ones who really need to sit down and do the math are developers running agents, coding workflows, and automated pipelines via the API. The 2.5x unit price hike is real, but so is the dramatic jump in token efficiency—meaning whether your actual bill goes up depends entirely on what your workload looks like. For long-running workflows with multi-turn tool calling that burn through mountains of tokens, switching over could very well end up cheaper. For single-turn, short tasks where you're merely using the model as a Q&A API, switching over is a pure 2.5x price increase with zero offsetting benefit. In that case, stick with Sol—it's still in the catalog, and still at its original price.
Finally, here's OpenAI's official announcement video if you want to see how they pitched it firsthand.
OpenAI's official launch video, "Introducing GPT-6 Astra." (Source: OpenAI Official YouTube Channel)
Reading this far, it might sound like I've been overly critical, so let me clarify: I'm actually optimistic about this release. The caveats above are about "not swallowing marketing figures whole," not "this model is bad." A 62.7% score under ARC's standard harness is still a massive leap over the previous generation's 7.8%, and beating the human median in action efficiency is genuinely impressive. Cutting the hallucination rate in half and dropping prompt injection success from 27% to 8.5% address two of my biggest everyday pain points; those two alone make it worth trying. As for the AGI claim, my attitude is to park the debate for now. Once it actually rolls out to my account, throwing my daily workflows at it to see whether it holds up in practice is far more productive than arguing over definitions here. Once I've spent more time with it, I'll be back with a full hands-on review.
Further Reading: Claude Fable 5.1 Launches: Why Your $20/Month Claude Pro Plan Can't Actually Run It, and How to Choose Between the Big Three AI Subscriptions, Gemini 3.8 Flash Launches: Official Benchmarks Tie or Beat Opus 5 Across Most Tasks, but Two Big Gaps Are Hard to Hide
Sponsored
Related

Apple's Sept 10 Event Is Official: With a Rumored NT$60,000 Foldable iPhone Ultra, Should You Wait or Buy Current Stock Now?
Apple officially confirms its Sept 10 event at 1:00 AM. The first foldable iPhone is rumored to start around NT$63,000 for 256GB—roughly 1.6 times the current iPhone 17 Pro. Here is our breakdown of how to watch the livestream, whether the first-gen foldable is worth the gamble, and how to pick from in-stock models if you're on a budget.

Major Chrome Zero-Day Under Active Exploit: Edge and Brave Users Need to Update Too
Google confirmed that CVE-2026-85046 in Chrome's V8 engine was actively exploited in the wild before a patch was released. Rated at CVSS 8.8, this marks the sixth exploited Chrome zero-day this year. Because the flaw resides in the Chromium engine itself, browsers sharing the same foundation—like Edge and Brave—are all affected, and the Android version requires a separate update. Here is a walkthrough of how to check and update each browser.

LINE Is Deleting Accounts by Year-End: Heads Up If You Only Linked Facebook Without a Phone Number!
The notices claiming "LINE will delete accounts without a linked phone number by year-end" are real, but they only target one specific group: accounts linked solely to Facebook with no registered phone number. Here is how to check if you're affected, how to set up your phone number, email, password, and Apple/Google link all in one go, the major pitfall where linking a number deactivates another account, and the only path left if you're already locked out.