GPT-6 Astra Lands at $10/$50 With Benchmark Claims to Check
A two-stage rollout, frontier-tier pricing, and a benchmark chart with at least one number that should make you suspicious.
Jump to 4 sections
OpenAI shipped GPT-6 Astra on September 3, 2026, opening it first to approved organizations before widening access the following day. It sits at the top of the GPT-6 line, above the Sol and Luna tiers that arrived over the summer, and it is priced accordingly.
The headline number for anyone building on it is $10 per million input tokens and $50 per million output. That is five times what Sol costs on output and a hundred times what Luna costs. Astra is not a drop-in upgrade for a production workload; it is a tier you route to deliberately.
What OpenAI says is different
The pitch is reasoning depth and agentic reliability rather than raw speed. OpenAI's published figures put Astra at the top of several evaluation suites, including FrontierMath's hardest tier and the ARC-AGI-3 abstraction benchmark, alongside a specialized security-testing variant.
Two of those claimed scores deserve a flag. A near-perfect result on an adversarial benchmark usually means the benchmark has been saturated, contaminated, or scoped narrowly enough that the number stops being informative. It does not automatically mean the model is weak — it means the number is not doing the work the marketing wants it to do.
Model Drop has not independently reproduced any of these figures. We are reporting what the vendor published, which is a different thing from confirming it, and we would treat every launch-day benchmark chart the same way. The methodology caveats in our guide to reading a model launch announcement apply directly here.
The pricing tier is the real story
Astra's cost structure tells you who it is for. At $50 per million output tokens, a single agent run that emits 40,000 tokens costs $2 before you count the input side. Multiply that across a support queue or a nightly batch job and the arithmetic stops working quickly.
| Tier | Input / 1M | Output / 1M | Sensible use |
|---|---|---|---|
| GPT-6 Astra | $10 | $50 | Hard reasoning, low volume |
| GPT-6 Sol | $2 | $10 | General production work |
| GPT-6 Luna | $0.10 | $0.50 | High-volume, simple tasks |
Those figures are OpenAI's list prices as of the launch window; check the OpenAI platform pricing documentation before you budget against them, because tiers and promotional rates move.
The three-tier structure itself is now the industry default. Anthropic runs the same shape — Anthropic's published Claude pricing lists Fable 5.1 at $10/$50, Opus 5.5 at $4/$20, and Sonnet 5 at $2/$10 — and Google splits Pro and Flash the same way. Every serious lab has concluded that one model at one price does not fit the market.
What this changes if you are shipping
Very little this week, and that is the honest answer. A frontier release at the top price tier is not something you migrate to on launch day. The sensible sequence is to benchmark it on your own evaluation set, price the delta against whatever you run now, and route only the traffic that genuinely needs the extra capability.
The workloads where a jump like this tends to pay for itself are narrow: multi-step agents that fail expensively, code changes against large unfamiliar repositories, and anything where a wrong answer costs more than $2 to clean up. Everything else belongs on a cheaper tier.
Independent evaluation will take a few weeks to arrive. Community leaderboards such as LMArena's public model rankings collect head-to-head human preference data over time, and third-party coding benchmarks like SWE-bench will post verified numbers on their own schedule. Those are worth more than any launch-day chart.
At Model Drop we will revisit Astra once there is independent data to compare against the vendor's own. Until then, the defensible summary is that OpenAI has a new top tier, it costs $10 and $50, and the capability claims are unconfirmed.
What to watch next
Three things. Whether the claimed reasoning gains hold up on tasks that were not in anyone's benchmark suite. Whether OpenAI cuts Astra's price within two quarters, as has happened with every previous top tier. And whether the specialized variants turn into a permanent product line or quietly disappear.
The release also lands inside an unusually dense launch window — four major labs shipped within days of each other, which we cover in our coverage of Anthropic's Fable 5.1 release two days earlier.
By Greg Halston, Staff Writer at Model Drop. Reported September 4, 2026; benchmark claims are the vendor's and have not been independently verified by Model Drop.
Model Drop covers AI launches — new models, platforms, features, and tools — for the people who have to decide what to actually ship on.