GPT-5.6 Price Cuts: OpenAI Slashes Luna and Terra Rates
Smart Market Insight Editorial
Editorial Team
This article may contain affiliate links. We only recommend tools we’ve personally tested. Read our full disclaimer.
OpenAI cut API prices on two of its three GPT-5.6 models on July 30, 2026 — Luna dropped 80%, Terra dropped 20% — and added a paid "Fast mode" for the flagship Sol model. None of the three models got smarter. The move is straightforwardly about cost, arriving one day before OpenAI announced it had crossed 1 billion active users across ChatGPT, Codex, and ChatGPT Work.
If you build anything on the OpenAI API — a support bot, a content pipeline, an internal agent — this is one of those quiet pricing updates that changes your monthly bill more than most product launches do. Here's exactly what changed, why OpenAI did it now, and how the new rates stack up against Claude and Gemini.
Quick Take
- What changed: GPT-5.6 Luna's API price fell 80%, from $1 / $6 to $0.20 / $1.20 per million input/output tokens. GPT-5.6 Terra fell 20%, from $2.50 / $15 to $2.00 / $12.00 per million tokens. Sol, the flagship, kept its price but gained an optional "Fast mode."
- Fast mode: Replaces OpenAI's old Priority Processing tier for Sol. It runs up to 2.5x faster than standard processing, at double the standard price ($10 input / $60 output per million tokens), with no change in output quality.
- Why now: OpenAI says efficiency gains in its models and serving infrastructure made the cuts possible; the announcement landed the same week OpenAI said it passed 1 billion active users and 2 million enterprise customers.
- Who benefits most: Anyone running high-volume, low-complexity workloads on Luna — classification, simple extraction, chat triage — gets the biggest relative savings. Sol users pay the same unless they opt into Fast mode.
What Actually Changed, Model by Model
GPT-5.6 ships in three tiers, and OpenAI only touched the cheap and mid-tier ones. According to OpenAI's own announcement, the changes take effect immediately for all API customers — no opt-in required, no code changes needed to get the lower rate.
| Model | Old price (in/out per 1M tokens) | New price (in/out per 1M tokens) | Change |
|---|---|---|---|
| GPT-5.6 Luna (smallest) | $1.00 / $6.00 | $0.20 / $1.20 | −80% |
| GPT-5.6 Terra (mid-tier) | $2.50 / $15.00 | $2.00 / $12.00 | −20% |
| GPT-5.6 Sol (flagship) | $5.00 / $30.00 | $5.00 / $30.00 (Fast mode: $10 / $60) | No change |
Pricing verified August 2026 from OpenAI's official announcement and OpenAI Developers' posted confirmation; confirm current rates on OpenAI's pricing page before budgeting, since API prices have moved more than once already this year.
Fast mode is the one genuinely new thing here rather than a straight discount. It's a rebrand and speed bump of what used to be called Priority Processing — any request already tagged priority automatically routes to Fast mode, so nothing breaks for existing customers. The pitch is simple: same Sol intelligence, up to 2.5x the throughput, for exactly double the per-token price. That's a real option for latency-sensitive products (voice agents, live chat) that were previously stuck choosing between Sol's quality and a faster-but-dumber model.
Why OpenAI Cut Prices Now
Two things happened in the same week that explain the timing. First, competition on price has genuinely intensified: Moonshot's Kimi K3 undercut GPT-5.6 Sol by roughly two-thirds on output pricing when it launched in mid-July, and DeepSeek's V4 lineup has kept pushing open-weight pricing down all year. Luna, OpenAI's cheapest model, was the one most exposed to that pressure — it's the tier most likely to get swapped out for a cheaper open-weight alternative on high-volume, low-stakes tasks.
Second, OpenAI said the cuts follow real efficiency work, not just a reaction to rivals: improvements to the models themselves and to the serving infrastructure they run on, which lowered the actual cost of running inference. That's a more sustainable story than pure margin-cutting, if it holds up — but it's also exactly what a company under competitive pressure would say regardless, so treat the "we got more efficient" framing as OpenAI's explanation rather than an independently verified fact.
The backdrop matters too. OpenAI said it passed 1 billion active users and 2 million enterprise customers around the same time, a milestone that took longer to reach than the company initially expected after topping 900 million weekly users back in February. At that kind of scale, a small per-token cut compounds into a large amount of retained volume — cheaper Luna pricing makes it less attractive for high-traffic, cost-sensitive customers to switch providers.
How the New Rates Compare to Claude and Gemini
Comparing "cheapest tier to cheapest tier" and "flagship to flagship" is the only fair way to read this, since the three vendors don't use identical naming:
| Tier | OpenAI | Anthropic | |
|---|---|---|---|
| Budget | GPT-5.6 Luna: $0.20 / $1.20 | Claude Haiku 4.5: $1.00 / $5.00 | Gemini 3.1 Flash-Lite: $0.30 / $2.50 |
| Mid-tier | GPT-5.6 Terra: $2.00 / $12.00 | — | Gemini 3.1 Pro: $2.00 / $12.00 |
| Flagship | GPT-5.6 Sol: $5.00 / $30.00 | Claude Opus 5: $5.00 / $25.00 | — |
All figures per million input/output tokens, verified against each vendor's published API pricing in early August 2026.
Luna's new price makes it the cheapest budget-tier model of the three by a wide margin — less than a fifth of Gemini's Flash-Lite output rate and a fraction of Claude Haiku's. At the flagship end, nothing changed: Claude Opus 5 still undercuts Sol on output tokens ($25 vs. $30 per million) while leading it on Anthropic's own agentic-coding benchmarks, so Sol's Fast mode is a speed play, not a price play, against Anthropic's current flagship.
Should You Actually Switch?
If you're already running high-volume tasks on Luna — chat classification, content tagging, simple summarization — this cut is close to free money; the same code, the same quality, at a fifth of the previous cost. Worth re-checking your usage dashboard this week rather than waiting for your next invoice to notice.
If you're on Terra for medium-complexity work, the 20% cut is welcome but not dramatic enough to justify migrating a stable Sol or Claude workload down a tier just to chase it — the quality gap between tiers is still real. And if your product genuinely needs lower latency more than lower cost, Fast mode is worth testing against your actual traffic before committing, since doubling your per-token spend to save a few hundred milliseconds only pays off for a specific class of live, user-facing product.
For a broader view of which AI subscriptions and APIs are worth paying for in 2026, see our full roundup of AI tools worth paying for, or browse the AI Tools category for more pricing and benchmark coverage as the field keeps moving.
Frequently Asked Questions
Did GPT-5.6 Sol get cheaper? No. Sol's standard pricing is unchanged at $5 input / $30 output per million tokens. The only new option is Fast mode, which costs double for up to 2.5x the speed.
Is Fast mode required to keep using Sol?
No, it's opt-in. Existing requests continue running at standard speed and price unless you explicitly request Fast mode, though requests already tagged priority route to it automatically.
How much cheaper is GPT-5.6 Luna now? Input tokens dropped from $1.00 to $0.20 per million (an 80% cut) and output tokens from $6.00 to $1.20 per million — the same 80% reduction applied to both sides.
Does this make GPT-5.6 cheaper than Claude or Gemini? At the budget tier, yes — Luna is now cheaper than both Claude Haiku 4.5 and Gemini's Flash-Lite tier. At the flagship tier, Sol is unchanged and still priced slightly above Claude Opus 5 on output tokens.
Why did OpenAI cut prices now instead of earlier? OpenAI points to efficiency gains in its models and infrastructure. The timing also lines up with intensifying price competition from open-weight models and OpenAI crossing 1 billion active users, both of which make holding volume at the low end more valuable.
Bottom Line
The GPT-5.6 price cuts are a real, verified change worth acting on if you use Luna or Terra in production — Luna's 80% cut in particular changes the math for high-volume, low-complexity workloads. Sol users get a new speed option, not a discount, and the flagship-tier competition with Claude Opus 5 hasn't shifted. Treat this as one more data point in a genuine industry-wide price war rather than a one-off promotion — compare it against how Claude, ChatGPT, and Gemini stack up overall before deciding where your workload actually belongs.
Related Articles
GPT-5.6-Cyber: Inside OpenAI's Gated Hacking AI
OpenAI's GPT-5.6-Cyber finds zero-days at a 95% success rate, but it's locked behind a vetted partner program. Here's what GPT-5.6-Cyber actually does and who can use it.
ChatGPT Atlas Shutdown: Why OpenAI Killed Its Own AI Browser
OpenAI shut down ChatGPT Atlas on August 9, 2026, folding its AI browser into ChatGPT itself. Here's why, and what actually replaces it.
ChatGPT Removes Free-Tier Chat Limits: What Changed, and Why
OpenAI is removing ChatGPT's free-tier text chat limits — here's what changed on August 6, why now, and how it compares to Claude and Gemini.
EU AI Labeling Rules Are Now Law: What It Means for AI Tools
EU AI Act Article 50 took effect August 2, 2026, forcing AI content labeling worldwide. Here's what it requires and who it actually affects.