OpenAI launched GPT-6 Sol and GPT-6 Luna yesterday, and every headline led with the same number: 50 percent off. GPT-6 Sol costs $2 per million input tokens and $10 per million output, down from $4 and $20 for GPT-5.6 Sol; GPT-6 Luna costs $0.10 and $0.50, down from $0.20 and $1.20, per VentureBeat's report. The halving is real and the models are generally available today on Amazon Bedrock and Microsoft Foundry. But the 50 percent is the least durable claim in the launch, and two quieter facts are the ones that will actually reshape how people architect agent stacks.
The launch, at list price
The family structure is now a clean three-tier ladder. GPT-6 Astra, launched earlier this month, anchors the top at $10 input and $50 output per million tokens (short context, Global Standard, per Microsoft's published pricing table). Sol is the mid-tier daily workhorse for coding, debugging and multistep workflows; Luna is the volume tier for extraction, summarization, routing and routine interaction. The step between tiers is almost exactly 10x on input: $10, $2, $0.10.
| Model | Input ($/MTok) | Cached input | Output ($/MTok) | vs predecessor |
|---|---|---|---|---|
| GPT-6 Astra | $10.00 | $1.00 | $50.00 | flagship, new family |
| GPT-6 Sol | $2.00 | $0.20 | $10.00 | 50% cheaper both ways |
| GPT-6 Luna | $0.10 | $0.01 | $0.50 | 50% in, 58% out |
| GPT-5.6 Sol | $4.00 | (varied) | $20.00 | predecessor |
Source: Microsoft Foundry and OpenAI published rates as collected by VentureBeat; OpenAI confirmed both halves on inputs and outputs across the family.
OpenAI's stated mechanism for the cut is not model cheapness for its own sake: improvements in caching and inference lowered serving cost, and the company says it is passing the savings through. On quality, the company claims roughly half as many factual mistakes as GPT-5.6 Sol on an internal factuality evaluation, which AWS's launch post repeats with the correct attribution: internal, OpenAI-run. Hold that label; it is the same caveat class we apply to every vendor eval in our benchmark field guide.
The claim nobody headlined: the cache tier is the real price
Buried in the pricing tables is the number that should reorganize agent budgets. Luna's cached input rate is $0.01 per million tokens: one tenth of its already-low input price, and one thousandth of Astra's standard input rate. Sol's cached input is $0.20, against $10 output. Both models support explicit prompt caching on Bedrock and Azure, which means marking instructions, tool definitions, policies and reference material once and paying roughly a tenth on every subsequent read.
Run the arithmetic on the shape of a real workload. A support agent with a 20,000-token system prompt (instructions, retrieval schema, policy text) answering 10,000 questions a day replays that prompt 10,000 times. Uncached on Luna that is 200 million input tokens a day, $20. Cached at $0.01, it is $2. Meanwhile the answers (output tokens) still bill at the full $0.50 rate, because generation is the part that still costs real compute, exactly the asymmetry we mapped in the Jev analysis and the Astra token-economics piece. The cache discount attacks the input side; the autoregressive tax on output survives untouched. A 100x discount on context reads and a flat price on generation tells you precisely where OpenAI's marginal cost still lives, and it is the same place it has always lived.
The corollary for builders: prompt caching stops being an optimization tip and becomes the pricing strategy. The right question is no longer which model but which tokens, because the effective blended rate of a cache-heavy pipeline on Luna can undercut the sticker price by an order of magnitude, and a cache-blind pipeline pays a 10x penalty for the same work.
Permanence is the competitive weapon
The second unheadlined fact is the one a spokesperson volunteered to VentureBeat: these are permanent prices, not introductory or promotional rates. That distinction exists because the price war is currently being fought with expiring coupons. The same VentureBeat pricing table shows Gemini 3.8 Flash sitting at $0.75/$3.75 through December 31, 2026, rising to $1.50/$7.50 on January 1, 2027. DeepSeek lists separate peak and off-peak rates. Google has run launch-window rates on its Flash line as standard practice.
Against that field, a permanent-pricing pledge is a different kind of claim: it is not "we are cheap," it is "you can write our number into your unit economics and not re-architect in January." For anyone building voice or agent pipelines where the model bill is a real line item, the difference between a promo price and a permanent price is the difference between a rounding error and a bet on someone else's marketing calendar. The same logic explains why Anthropic made Claude Sonnet 5's introductory $2/$10 permanent in August: at the mid tier, price certainty has become the tiebreaker once price level converges, because Sol and Sonnet 5 now sit at identical rates, dollar for dollar, input and output.
There is a credible skeptic's reading, and the article should carry it. "Permanent" is a statement about intent with no contractual force outside enterprise commits; the same industry that cut prices 50 percent can quietly raise them when the competitive pressure lifts, and Google's own January price increase will be the next natural experiment in how long these pledges survive. The permanent-pricing claim is honest as written and unverifiable as a promise. What makes it still valuable is that the direction has structural support: serving costs really are falling, per the cost chain we traced, so holding a low price does not require charity, it requires not grabbing margin that the efficiency gains hand you. Whether OpenAI refrains from grabbing it is the actual thing being pledged.
Where Luna sits in the price war, and where it does not
Luna's $0.10/$0.50 lands it near the bottom of the mainstream API field on sticker price, with only a handful of Chinese-lab models undercutting it (Meta's contributor-tier Spark at $0.10/$0.20, Xiaomi's MiMo Flash at $0.14/$0.28, DeepSeek's off-peak Flash at $0.15/$0.60, per the same VentureBeat table). But the honest comparison has three asterisks. First, those Chinese rates are frequently variable (peak/off-peak, limited-time promos), so Luna's flatness is a feature they lack. Second, Luna is a closed frontier-lab model at an open-weights price point, which is the actual industry story: frontier API pricing is converging toward the cost floor that open-weight hosts established, and our local-versus-API break-even math showed that convergence is precisely what pushes self-hosting break-evens further out for most teams. Third, at $0.50 per million output tokens Luna's decision-shaped calls approach the price class where specialized decision models like Jev's $0.0004 per decision compete: a Luna call doing a routing job costs well under a cent, so the Jevons floor is no longer exotic hardware economics, it is a mainstream OpenAI price tier.
The 58.3 percent output cut on Luna also deserves note because it breaks the family's tidy 5x output-to-input ratio in the buyer's favor, while Sol and Astra keep the 5x exactly. Output discounts are the ones that change agent economics, since agentic pipelines spend most of their budget generating, and OpenAI cutting output faster than input on the volume tier is a small tell about which workload it most wants to capture: the high-frequency machine one, not the human chat one.
Our Read
Three labeled opinions. First, the durable innovation in this launch is the cache tier, not the halving: a 10x discount on cached reads makes context architecture the largest cost lever in most agent systems, and it quietly rewards the boring engineering (stable prompts, marked reusable blocks, disciplined schemas) that generic "which model is smarter" coverage never prices in. Second, identical pricing at the mid tier across two frontier labs (Sol and Sonnet 5, $2/$10 to the dollar) is the cleanest evidence yet that at this capability level the market price is a cost price, meaning the next differentiation will come from serving efficiency and reliability, not model intelligence, which is exactly the specialization logic our cost-chain piece predicted. Third, treat both halves of the quality story (half the factual mistakes, better task completion) as vendor claims until independent evals land; the pricing facts are checkable today on three platforms, the accuracy facts are internal OpenAI numbers, and a launch whose price is verifiable and whose quality is not should change your evaluation plan accordingly: run your own evals, because the price cut alone will not tell you whether the cheaper model still clears your quality floor.
Outlook
Watch three things. The January 1, 2027 Google price increase is the scheduled test of whether the industry's promo pricing was price discovery or price marketing; if Google raises and OpenAI holds, permanence becomes the standard competitive axis. Second, watch whether cached-input rates (not standard rates) become the headline number in the next round of launches, because once buyers start comparing effective blended rates the sticker prices stop meaning much. Third, watch for the first major lab to price output tokens at zero for machine-only tiers, which would complete the journey our cost chain described from $30,000 wafers to free outputs, this time at the biggest API in the world. The 50 percent headline will be old news within a quarter. The one-cent cache and the permanence pledge are the parts that compound.