OpenAI Just Slashed GPT-5.6 Luna Prices by 80 Percent: the AI Price War Is Here

OpenAI cut GPT-5.6 Luna prices by 80% and Terra by 20%. Luna now costs $1.40 per million tokens, putting it in direct competition with DeepSeek and Google.

OpenAI Just Slashed GPT-5.6 Luna Prices by 80 Percent: the AI Price War Is Here

The AI model market just shifted from a capability arms race to a price war. And OpenAI just fired the biggest salvo yet.

Today the company slashed the cost of GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20%, while leaving its flagship Sol unchanged. Luna, the smallest and fastest model in the GPT-5.6 family, now costs $0.20 per million input tokens and $1.20 per million output tokens, for a combined $1.40 per million tokens. Terra drops to $2 per million input and $12 per million output, matching Google’s Gemini 3.1 Pro Preview pricing at $14 combined.

This isn’t a minor promotional discount. It’s a structural repricing that brings an OpenAI frontier-series model into direct competition with the low-cost inference tier that Chinese providers like DeepSeek and Xiaomi have dominated for months.

Sam Altman announced the move on X with characteristic brevity: “major price cuts today.”

Luna enters the budget tier

When OpenAI launched the GPT-5.6 series three weeks ago, Luna was priced at $7 per million combined tokens. At $1.40, it now costs less than Google’s Gemini 3.5 Flash-Lite ($2.80) and far less than Gemini 3.6 Flash ($9). It undercuts OpenAI’s own GPT-5.4 by roughly 13x.

It’s not the absolute cheapest model on the market. DeepSeek’s flash model costs $0.42 per million combined tokens. Xiaomi’s MiMo-V2.5 Flash is at $0.40. But Luna brings frontier-level capability, including tool use, multi-step workflows, and the ability to drive agentic loops, into a price range that makes high-volume deployment economical in a way it simply wasn’t before.

Blitzy CTO Sid Pardeshi put it in practical terms: Luna moved his team from a single structured-output call to a full tool-calling agent loop, increased prompt-cache reuse from 24% to 90%, and delivered an 87% cost reduction versus GPT-5.4 mini. Replit’s Head of AI called Luna “the closest we’ve come to intelligence too cheap to meter.”

Terra matches the mid-range

The 20% Terra cut to $14 combined is less dramatic on its face, but context matters. Terra now sits at the exact same combined price as Google’s Gemini 3.1 Pro Preview, $14 per million tokens, and offers comparable intelligence for everyday work. Notion’s AI team said Terra delivers “comparable quality to GPT-5.5 at half the cost per task and in 60% less time.”

The Terra pricing also puts pressure on Anthropic. Claude Opus 5, which Anthropic positioned as the affordable-but-capable option, costs $30 combined ($5 input, $25 output), more than double Terra’s price. And Claude Sonnet 4.6, the mid-tier Anthropic model, runs $3/$15, which is more expensive than Terra on input while roughly matching on output.

Why the cuts happened now

The official story is efficiency gains. OpenAI published a detailed post yesterday explaining how GPT-5.6 Sol itself helped optimize the serving stack, rewriting production kernels, designing experiments to improve token generation, and monitoring training runs. The company says these improvements reduced end-to-end serving cost by 20% and increased token-generation efficiency by over 15%.

But the timing is not accidental. Google launched Gemini 3.6 Flash and 3.5 Flash-Lite nine days ago, both built around low inference costs for agent workloads. Anthropic shipped Claude Opus 5 at the same $30 combined price as Opus 4.8 but with near-Fable-5 capability. And Chinese open-weight models from DeepSeek, Moonshot AI, and others continue to pull pricing downward while closing the capability gap.

A CNBC report on the cuts framed it as OpenAI facing pressure to cater to a “more cost-sensitive customer base” while fending off Chinese competition. That’s accurate, but it understates the structural shift. This is a market that is rapidly commoditizing model intelligence at the low and mid tiers, and OpenAI is choosing to lead on price rather than defend a premium position.

Sol gets faster, not cheaper

The flagship Sol is unchanged on pricing at $5/$30 per million tokens. But it gained a new Fast mode that delivers up to 2.5x throughput at 2x the price ($10/$60). Fast mode replaces the old Priority Processing tier and is backward compatible.

This is a smart play. For the workloads that genuinely need frontier intelligence, complex reasoning, planning, and high-stakes analysis, latency matters more than marginal token cost. Sol Fast gives teams an escape valve without forcing them to cobble together their own routing logic.

What this means for developers and businesses

If you’re building on the OpenAI API, the calculus just changed. Luna is now cheap enough to use as a default for high-volume agent loops, classification pipelines, and background automation. Terra is the sensible choice for everyday knowledge work. Sol stays in its lane as the reasoning engine you call when the answer actually matters.

The broader take is that model pricing is in a deflationary spiral, and that’s good for anyone building on these APIs. Each round of efficiency gains, from better architectures and better serving infrastructure to better routing, gets passed down as lower prices rather than captured as margin. That pattern has held for every major model release cycle in the last two years, and the GPT-5.6 cuts suggest it’s accelerating.

The companies that will benefit most: anyone running high-volume agent workloads, customer-facing AI features at scale, or internal automation pipelines where every millisecond and every token hits the budget. The companies that should be nervous: anyone whose moat was “we have a cheaper model” rather than “we solve a specific problem better than anyone.”

The competitive landscape after today

The VentureBeat pricing comparison table tells the story. Luna at $1.40 sits between DeepSeek’s $0.42 and Google’s $2.80 in the combined-price ranking. Terra at $14 anchors the mid-range. Sol at $35 holds the frontier alongside Claude Opus 5.

But raw price-per-token is only part of the equation. The GPT-5.6 models bring tool-use, structured output, and agentic capabilities that the cheaper Chinese flash models don’t always match. The question for developers is whether the capability gap justifies the price gap, and as Luna pricing shows, that gap is narrowing fast.

OpenAI’s strategy seems clear: use the efficiency flywheel (better models optimize the infrastructure, which funds better models) to compete on both capability and cost simultaneously. Google is running the same play. Anthropic is betting that enterprises will pay a premium for safety and reliability. The Chinese labs are betting that open-weight distribution and aggressive pricing will win the volume game.

I’m not sure any of them are wrong. But the winner of this race will be the one that makes intelligence cheap enough that the question isn’t “can we afford to use AI for this?” but “can we afford not to?”

For context on the GPT-5.6 family launch, here is my earlier coverage of the announcement. And if you are trying to decide which model to build on, my AI coding agent comparison covers how these models perform in real development workflows.

Tony Simons

Reviewed & Written By

Tony Simons

Independent tech reviewer and creator of Tony Reviews Things. 14 years of hands-on testing, software auditing, and workflow automation. I test the gear so you don't waste your money on junk.

Submit a Take

Your email address will not be published. Required fields are marked *