GPT-6 Astra Is Out at Fable Prices: The New Agentic Coding Cost Math
OpenAI launched GPT-6 Astra at $10/$50 per million tokens. It undercuts Claude Fable on coding cost per task, yet costs more than Sol on intelligence work.
On September 3, 2026 OpenAI launched GPT-6 Astra, its new flagship reasoning and agentic model, and the pricing alone forces a re-baseline for anyone carrying a heavy API bill. Astra lists at $10 per million input tokens and $50 per million output, exactly the rate card of Claude Fable 5 and Fable 5.1, and 2.5 times the price of GPT-5.6 Sol at $4/$20. On paper that looks like OpenAI matching Anthropic at the top of the market. Under the hood the interesting story is cost per completed task, because that is where Astra’s token efficiency changes the math in ways a headline price number hides.
What OpenAI Actually Released
Astra is the first model OpenAI describes as built on its Stargate compute and the first it calls “lightly looped,” a small looped-reasoning step that lets the model reuse internal state across a long agent run rather than restart from scratch. It is rolling out to a limited set of organizations immediately, with ChatGPT Plus, Pro, Business, and Enterprise users, plus the OpenAI API and AWS, following over the coming days. The API model label is gpt-6-astra.
The capability specs are genuinely frontier level. Astra saturates ARC-AGI-3 at 99.9 percent and FrontierMath at 97.6 percent, hits 100 percent on ExploitBench against Sol’s 78.5 percent, and holds 96.3 percent accuracy out to 1M tokens on long-context tests. OpenAI frames it as the arrival of the automated AI engineer: a model that can pick and train models, keep pipelines saturated, deploy an entire system in one shot, and command fleets of subagents. Independent reviewers who burned tens of billions of tokens on early access confirmed those claims for real agentic work.
The Two Numbers That Matter for Your Bill
List price tells you the ceiling per token. It does not tell you what you actually pay per finished task, and that second number is where Astra and its rivals genuinely diverge. Independent benchmarking from Artificial Analysis splits the picture cleanly in two.
On the Artificial Analysis Coding Agent Index, Astra scores 67 in the Codex harness, roughly equal to Claude Opus 5 and Fable 5 and just behind Fable 5.1 at 70. The important part is token efficiency. Astra uses about one third of the tokens that GPT-5.6 Sol does at max effort in the Codex harness, and roughly one fifth of the tokens of Claude Opus 5. At max effort, Astra costs about the same per coding task as Sol while scoring two points higher, and it costs less than half of what Claude Fable 5 costs per task for the same score. For heavy users running long agentic coding sessions, that is the headline: parity-list-price model, dramatically better cost per completed task.

On the Artificial Analysis Intelligence Index the story flips. Astra scores 61 at max effort, equal to GPT-5.6 Sol and five points below Claude Fable 5.1. It does cut output tokens about 10 percent versus Sol at max effort, but because the price is 2.5 times higher, Astra ends up about 75 percent more expensive per task than Sol on the intelligence-type work that index measures. Translation for a heavy user: if most of your spend is agentic coding, Astra is likely a cost-per-task win. If a meaningful share of your bill is knowledge work, research, and agentic writing where Sol was your baseline, moving those to Astra at max effort makes them more expensive, not less.
Cache and Context Change the Leverage
Both of Astra’s dimensions reward the same discipline that already pays on Fable. Cache reads carry the standard 90 percent discount off input, and cache writes carry a 25 percent premium. With long-context performance holding at 96 percent out to 1M tokens, the economic case for holding a large cached working context on Astra is strong: the tokens you read back from cache cost a tenth of a fresh input, and Astra is one of the few models that keeps working correctly at that scale.
Independent testing found Astra sustained about 33 tokens per second while managing agent fleets, which puts its all-in scaled cost around $6 an hour. That is not a guaranteed rate, it depends entirely on your concurrency and effort level, but it is the right mental frame: for output-heavy agentic work, Astra behaves like a very cheap engineering hire rather than a per-token meter. Dial it to the highest effort with dozens of parallel subagents and you will burn far more than $6 an hour, and the model is good enough at parallelizing that it will happily help you do so.

What Heavy Users Should Do Now
Do not re-route everything the day access lands on your account. The rollout starts with a limited set of organizations, so most subscribers and API users do not have Astra yet. Use the gap to prepare.
Measure cost per task, not cost per token. If you track spend against Claude Fable or GPT-5.6 Sol, re-run your own benchmark suite against Astra on the codebase you ship. The Artificial Analysis numbers are directionally reliable, but your cache-hit ratio, agent depth, and prompt scaffolding will shift the result. Do not bill clients on a headline $10/$50 rate; bill on measured per-task cost.
Split your routing by task family. Agentic coding looks like a price win on Astra. Long-horizon intelligence and knowledge work at max effort looks roughly 75 percent pricier per task than Sol. The efficient stack is Astra for the coding surface and Sol, Luna, or a capable open-weight model for pure knowledge work, with a router deciding per request. A single default model for everything now leaves real money on the table in one direction or the other.
Lean into cache and long context, but meter the writes. Cache reads at $1 per million are the cheapest tokens you will buy this quarter, and Astra holds quality at 512K to 1M tokens. The 25 percent premium on cache writes means you want to pay the setup cost once and reuse heavily. If your workflow is many short stateless calls, you get almost none of Astra’s advantage; if it is long persistent agent sessions, you get most of it.
Re-verify token-efficiency claims with your own telemetry. The gap between list price and effective cost is wider for Astra than for almost any recent flagship because its token efficiency is the actual product. Log token counts per task on Astra versus your current default before you commit. Cheap per task only shows up in your own usage data.
The Net for Heavy AI Users
GPT-6 Astra is the cleanest example yet of why the AI price war stopped being about dollars per million tokens and became dollars per completed task. At the identical $10/$50 rate as Claude Fable, Astra undercuts Fable by more than half on measured coding cost per task, collapses token use to a third of GPT-5.6 Sol, and yet costs 75 percent more than Sol on the intelligence work most people run their dashboards on. The same price number buys very different outcomes depending on the workload. For heavy users the durable move is not to chase the launch. It is to instrument cost per task, split routing by workload, and let your own usage data, not the press release, decide where Astra earns its place in your stack.
Now available
Stop guessing your AI limits
The Mac app and web dashboard watch your Claude, ChatGPT, Gemini and more, and warn you before quotas hit.