OpenAI just made its cheapest frontier model dramatically more affordable — and the timing is anything but accidental. The company slashed prices on two models in its GPT-5.6 series, cutting GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20%, while simultaneously launching a premium Fast mode for its flagship GPT-5.6 Sol. CEO Sam Altman announced the changes publicly on X, calling them “major price cuts today.” The moves land roughly three weeks after the GPT-5.6 series first became broadly available — and just days after rivals Google and Anthropic each made their own cost-efficiency plays.
Summary
Key takeaways
- GPT-5.6 Luna’s combined token price dropped 80% to $1.40 per million tokens, undercutting Google’s Gemini 3.5 Flash-Lite ($2.80) and Gemini 3.6 Flash ($9).
- GPT-5.6 Terra fell 20% from $17.50 to $14 per million tokens combined, matching Google’s Gemini 3.1 Pro Preview for context windows up to 200,000 tokens.
- Sol Fast mode adds up to 2.5x throughput at $70 per million tokens combined — double the Standard rate — for latency-sensitive production workloads.
- Anthropic’s Claude Opus 5 remains priced at $30 combined per million tokens, delivering near-Fable 5 performance at the same rate as Opus 4.8.
- OpenAI’s Luna now competes directly in the low-cost model tier alongside offerings from Google, Xiaomi, DeepSeek, and MiniMax.
OpenAI Announces Major Price Cuts Across the GPT-5.6 Series
The numbers tell a stark story. Luna, the smallest and fastest model in the GPT-5.6 lineup, previously carried a combined input-and-output price of $7 per million tokens. After the cut, that figure collapses to $0.20 per million input tokens and $1.20 per million output tokens — a combined $1.40. For high-volume applications processing millions of requests daily, that gap is enormous.
80% Price Reduction for Luna
At $1.40 combined per million tokens, Luna now sits below Google’s Gemini 3.5 Flash-Lite at $2.80, and far below Gemini 3.6 Flash at $9. It also undercuts OpenAI’s own GPT-5.4. Luna does not claim the absolute lowest token price in the market — models from Xiaomi, DeepSeek, and MiniMax still undercut it on raw cost — but it marks the first time an OpenAI frontier-series model has entered that competitive pricing band.
That repositioning matters beyond the numbers. Luna targets high-throughput, low-latency workloads — summarization, classification, routing, lightweight real-time assistants — where the cost per individual request compounds at scale. Moving into this tier means OpenAI is now competing directly for workload categories it previously ceded to smaller, cheaper models.
20% Price Reduction for Terra
Terra’s cut is more modest but strategically pointed. The combined token price dropped from $17.50 to $14 per million tokens, matching Google’s Gemini 3.1 Pro Preview for context windows of 200,000 tokens or less. As Krea AI’s Nic Dunz noted on X, Terra also now undercuts OpenAI’s own GPT-5.4, which remains priced at $2.50 per million input and $15 per million output tokens — making Terra the better value at roughly one-thirteenth the cost on a per-intelligence basis. Terra is designed for general production deployments where capability and efficiency need to be balanced, not maximized in one direction.
Introduction of Sol Fast Premium Mode
Sol moves in the opposite direction. Standard pricing remains at $5 per million input tokens and $30 per million output tokens. The new Sol Fast mode charges $10 per million input tokens and $60 per million output tokens — a combined $70 — delivering up to 2.5 times the throughput without altering the underlying model’s intelligence. Rather than lowering the price of its most capable tier, OpenAI is charging a premium for latency advantages, signaling that for complex reasoning and agentic workloads, speed has its own market.
Competitive Positioning of OpenAI’s GPT-5.6 Models
The three GPT-5.6 tiers now map onto clearly distinct market segments. Luna competes in the low-cost inference market. Terra targets the mid-market pro tier. Sol anchors the frontier reasoning category.
Luna Competes in the Low-Cost AI Segment
The 80% Luna cut transforms OpenAI’s competitive footprint. Previously, the GPT-5.6 series was largely a premium offering. Now one tier sits inside the same pricing bracket as models from Google, Xiaomi, DeepSeek, and MiniMax. According to third-party analysis from Artificial Analysis, Luna outperforms Gemini 3.6 Flash and Gemini 3.1 Pro in intelligence benchmarks — meaning Luna’s cost-per-intelligence ratio has shifted materially in OpenAI’s favor. As AI startup Cognition noted on X, GPT-5.6 now “sits on the pareto curve of price/performance efficiency,” offering among the most favorable intelligence-to-cost ratios on the market.
Terra Matches Google’s Gemini 3.1 Pro Pricing
Terra’s $14 combined price point creates a direct match with Gemini 3.1 Pro Preview for mid-range context workloads. The wider gap it creates within OpenAI’s own lineup is notable too: Luna now costs one-tenth of Terra on a simple combined-token basis, while Terra costs 60% less than Sol Standard. The three tiers are no longer closely spaced — they represent genuinely different price-performance trade-offs.
Sol Targets Complex Reasoning Workloads
Sol’s positioning has not changed. It remains the model for advanced coding, multi-step planning, and tool-using agentic systems — workloads where reasoning depth justifies higher per-token costs. The Fast mode addition extends Sol’s appeal to enterprises that need frontier intelligence but cannot absorb the latency of Standard throughput. At $70 combined per million tokens, Sol Fast is the most expensive configuration in the lineup, a deliberate premium for time-critical deployments.
Market and Industry Implications of the Price Cuts
OpenAI’s timing reflects pressure from multiple directions simultaneously. According to reporting by CNBC, enterprises have grown increasingly cost-sensitive, scrutinizing AI bills that have at times reached billions of dollars. The era of unlimited AI usage without tracking costs — what some described as “tokenmaxxing” — has given way to a sharper focus on return on investment. That shift is forcing all frontier model providers to rethink what they charge and why.
Shift from Model Access to Production Economics
What OpenAI described in its release as a focus on “advancing both capability and efficiency so each generation of intelligence can accomplish more work at a lower cost” reflects something broader than a routine pricing adjustment. The competition among frontier providers has moved beyond which model is most capable. The question now is which provider offers the most predictable, cost-efficient path to running AI at production scale. OpenAI’s cuts are a direct response to that shift, and they reframe the GPT-5.6 series from a premium access product into a cost-competitive deployment platform.
Comparative Strategies of OpenAI, Google, and Anthropic
Each of the three major frontier providers has chosen a different mechanism to lower the total cost of production AI. OpenAI is directly cutting per-token rates. Google, with Gemini 3.6 Flash and Gemini 3.5 Flash-Lite, is pairing lower prices with reductions in token consumption and tool calls — Gemini 3.6 Flash reportedly uses 17% fewer output tokens than Gemini 3.5 Flash on the Artificial Analysis Index, with savings reaching up to 65% on some long-horizon engineering tasks. Anthropic took a third path: Claude Opus 5 costs the same $30 combined per million tokens as Opus 4.8, but delivers near-Fable 5 performance — effectively lowering the price per unit of capability without moving the sticker price. Anthropic also added an adjustable effort setting allowing developers to trade reasoning depth for speed and token savings.
All three approaches target the same operational metric: the total cost of completing real production work, not just the advertised price of a single token. For enterprise buyers evaluating which platform to standardize on, that distinction now drives the conversation more than raw benchmark scores.
What remains unresolved is whether these price cuts are sustainable at scale or represent a land-grab moment driven by competitive pressure. Chinese open-weight models, including offerings from DeepSeek and MiniMax, continue to push the floor lower on pure token cost, while closed-model providers invest in higher infrastructure expenditure. Amazon, for its part, raised its 2026 capital expenditure to $220 billion, per CNBC, partly driven by rising memory costs — a reminder that the economics of producing cheap AI tokens are still under pressure at the infrastructure level. For now, OpenAI is betting that frontier-quality intelligence at low-cost pricing is a combination the market will pay for, even if not everyone at the bottom of the price table can match it.
FAQ
How much did OpenAI reduce the price of GPT-5.6 Luna?
OpenAI cut the price of GPT-5.6 Luna by 80%, lowering the combined input and output token cost to $1.40 per million tokens — down from a previous combined price of $7 per million tokens.
What is the new pricing strategy for OpenAI’s GPT-5.6 Sol model?
OpenAI added a premium Sol Fast mode that offers up to 2.5 times throughput at double the cost of the Standard mode, charging $70 per million tokens combined ($10 input, $60 output). Sol Standard pricing remains unchanged at $35 combined per million tokens.
How does OpenAI’s Luna pricing compare to Google’s Gemini AI models?
Luna’s $1.40 combined token price is cheaper than Google’s Gemini 3.5 Flash-Lite at $2.80 and significantly below Gemini 3.6 Flash at $9 per million tokens combined.
What workloads are the GPT-5.6 Luna, Terra, and Sol models designed for?
Luna targets high-throughput, low-latency tasks such as summarization, classification, and routing. Terra balances capability and efficiency for general production workloads. Sol focuses on complex reasoning, advanced coding, multi-step planning, and agentic systems.
Article produced with the assistance of artificial intelligence and reviewed by the editorial team.

