tokenkarma is in beta. Expect rough edges, and your feedback shapes what we fix next.
7 min read B2C power user

OpenAI Cuts GPT-5.6 Terra and Luna Prices: What Heavy AI Users Save

OpenAI slashed GPT-5.6 Terra prices by 20% and Luna by 80% on July 30. Here is exactly what heavy AI users save per million tokens and how to rebalance your spend.

OpenAI Cuts GPT-5.6 Terra and Luna Prices: What Heavy AI Users Save

OpenAI cut the price of its GPT-5.6 Terra and Luna models on July 30, 2026, roughly three weeks after their public release. Terra dropped 20% and Luna dropped 80%. For heavy AI users watching a five-figure monthly API bill, this is the biggest single line-item reduction in recent memory.

At the new prices, Luna now costs $0.20 per million input tokens and $1.20 per million output tokens. That puts OpenAI’s fastest model squarely into the budget tier that Chinese open-weight models had claimed for themselves. Terra, the mid-tier workhorse, now sits at $2 per million input and $12 per million output.

The flagship Sol model is unchanged at $5/$30 per MTok.

Why OpenAI cut prices now

OpenAI is under real pressure from three directions at once. Enterprises have grown sharply cost-sensitive, and many refuse to deploy expensive models without a clear return on investment. Chinese open-weight startups, led by Moonshot AI’s Kimi K3, are shipping models that match or beat US flagships on several benchmarks at a fraction of the API cost. And Anthropic, Google, and Microsoft are all advertising cheaper alternatives in the same week.

Anthropic released Claude Opus 5 on July 24 at $5/$25 per MTok, matching Fable 5 performance on key benchmarks at half the price. Google launched Gemini 3.6 Flash to undercut Chinese models on per-task cost. Microsoft touted a cheap-but-performant cybersecurity model on its earnings call. OpenAI’s price cut is a direct response to this crowding.

The move is landing on a strong base. OpenAI’s July revenue topped its entire second quarter, driven by the GPT-5.6 release. The price cuts are a bet that lower unit prices will hold volume high enough to keep total revenue climbing.

What the new per-token math looks like

Pricing comparison chart showing Terra and Luna price reductions

ModelOld inputNew inputOld outputNew outputCut
GPT-5.6 Sol$5$5$30$30None
GPT-5.6 Terra$2.50$2$15$1220%
GPT-5.6 Luna$1$0.20$6$1.2080%

For a concrete picture, take a workload of 100 million input tokens and 20 million output tokens per month, a realistic profile for a team running heavy coding and agent workflows.

On Luna, the monthly bill drops from $220 to $44. That is an 80% reduction, and it changes the tier economics entirely. Luna was the fast, cheap option for high-volume calls. Now it is dramatically cheaper than most of the market.

On Terra, the same workload drops from $550 to $440, a 20% reduction. For teams running long agent sessions where the output stream dominates, the output cut from $15 to $12 per MTok is where most of the savings land.

What this does to the cross-provider comparison

Comparative bar chart of OpenAI, Anthropic, and Google model cost per million tokens

The price cuts reshape the tier-vs-tier comparison for heavy users. GPT-5.6 Luna at $0.20/$1.20 now undercuts Claude Sonnet 5’s introductory pricing ($2/$10 per MTok) by an order of magnitude on input, and it is competitive with or below several Chinese open-weight API offerings.

For context on the low end of the market:

  • GPT-5.6 Luna: $0.20 input / $1.20 output per MTok
  • DeepSeek V4 Pro: $0.435 / $1.68 per MTok (via standard API pricing)
  • Gemini 3.5 Flash-Lite: $0.30 / $2.50 per MTok
  • Kimi K3 open weights through Telnyx: $2.70 / $13.50 per MTok

Luna is now among the cheapest frontier-adjacent APIs available from a major US provider. That matters because OpenAI can afford to make Luna aggressively cheap: it is the fastest model of the three, built for high-throughput, low-latency responses where token volume is the cost driver but each request is lighter.

Where the real savings are for heavy users

The temptation is to move everything to Luna and call it a day. Resist it. Luna is optimized for speed, not for the hardest reasoning tasks. For agent orchestration, classification, extraction, high-frequency tool calls, and routine coding, Luna at $0.20 input is a gift. Put the big reasoning workloads, multi-step planning, and prompt-sensitive benchmarks on Terra at the new $2/$12 rates.

The output-token economics are worth a second look. Most agent workloads are output-heavy: thinking tokens, tool-call stacks, code generation, and long rewrites all land on the output side. Luna’s output cut from $6 to $1.20 is the single most impactful number in this announcement. On short inputs and long outputs, the ratio swing is extreme.

If you are on the OpenAI API today, re-audit which tier each traffic pattern belongs on. A team routing 60% of its volume to Sol out of habit can often move that share to Terra or Luna with no quality loss on the lighter tasks, and the bill falls hard. Rewriting a single routing rule is the cheapest optimization most heavy users will ship this month.

What heavy AI users should do now

  1. Re-scope your tier assignments. Run a two-day token audit and classify traffic by reasoning demand. Anything that does not need Sol should move to Terra or Luna.
  2. Re-shop the low end. At $0.20 input, Luna competes with open-weight APIs on price. If your only constraint was cost, retest Luna against your cheap-tier provider before renewing contracts.
  3. Watch for the follow-on cut. OpenAI reacting three weeks after launch signals an aggressive price war. Anthropic and Google are likely to respond. Do not lock multi-year enterprise agreements on today’s list prices.
  4. Guard the reliability assumption. Deep price cuts on a flagship series can precede capacity refocusing. If your team depends on Luna at scale, monitor error rates and latency in the first two weeks.

OpenAI just made the low end of its lineup dramatically cheaper, and that is a direct win for heavy AI spenders. The teams that rebalance their tier mix this week will capture most of the savings. The teams that leave every workflow on Sol will keep paying for intelligence they are not using.

For teams tracking spend across Terra, Luna, and the fast-moving competition, the price cut is the reminder that API economics shift monthly. The winning posture is a routine audit of what you run, where you run it, and whether the frontier tier is actually earning its premium.