memujo
AI7 min read

GPT-6.1 Sol Holds Price and Cuts Cost Per Task

OpenAI's GPT-6.1 Sol keeps GPT-6 Sol's exact sticker prices yet claims near-Astra results at a fraction of the cost per task. Here is the arithmetic.

By Alice

In this article
  1. 01The launch in context
  2. 02What the documentation actually says
  3. 03The benchmark claims, with the vendor label attached
  4. 04What This Means: the arithmetic behind a flat price
  5. 05The Ultrafast question: who is actually serving these tokens
  6. 06Outlook

OpenAI launched GPT-6.1 Sol at DevDay on September 29, and the most interesting thing about it is a number that did not move. The model costs exactly what GPT-6 Sol costs: $2 per million input tokens and $10 per million output, per the official model page. What changed is everything behind the sticker. OpenAI claims the model lands near its flagship GPT-6 Astra on agentic benchmarks at roughly one-fifth the cost per task, which means the price of a unit of capability, not a unit of tokens, is where the deflation happened.

The launch in context

DevDay 2026 brought more than 20 announcements, including Dots (always-on agents powered by GPT-6 Astra), ChatGPT Space, and an Ultrafast speed tier across the API and Codex, as IT Voice summarized. GPT-6.1 Sol is the model aimed at builders who want most of Astra without paying flagship rates on every call.

The launch also had an unusual prelude. One day earlier, OpenAI announced it had abandoned GPT-6.1 Astra because the model frequently ignored instructions, as AFP reported through theSun. That same report puts the commercial pressure in perspective: Anthropic out-earned OpenAI in the second quarter, $11.6 billion against $6.7 billion per the Wall Street Journal, while OpenAI's Q2 operating loss widened to $12.3 billion. Shipping a mid-tier model that approaches last cycle's flagship economics is what a company does when per-task cost, not raw capability, is the contested axis.

What the documentation actually says

The spec sheet tells you this is an Astra-class model wearing a Sol price tag. Both models share identical window parameters, and the difference shows up in what is switched off.

Parameter GPT-6.1 Sol GPT-6 Sol GPT-6 Astra
Input, per 1M tokens $2 $2 $10
Cached input, per 1M $0.10 $0.20 $1
Cache write, per 1M $2.50 $2.50 $12.50
Output, per 1M tokens $10 $10 $50
Context window 1,050,000 1,050,000 1,050,000
Max output tokens 128,000 128,000 128,000
Knowledge cutoff Apr 30, 2026 Apr 20, 2026 Apr 30, 2026
Reasoning efforts low to max none to max low to max

Pricing figures are from the GPT-6.1 Sol and GPT-6 Astra docs pages. Two details in that table are easy to miss and both matter for agent workloads.

First, cached input dropped from 10% of the base rate to 5%, so $0.10 per million versus $0.20. Prompts over 272K input tokens trigger a 2x multiplier on input and 1.5x on output for the whole request, identical to Astra's long-context surcharge. A mid-tier model now inherits flagship long-context behavior, including the flagship penalty for abusing it.

Second, GPT-6.1 Sol drops support for none and minimal reasoning efforts, matching Astra. The migration guide is explicit: teams running GPT-6 Sol at none for cheap, fast calls need to start at low and re-evaluate. There is no zero-thinking mode here. Every call does some reasoning, which shifts cost from "did you pay for intelligence" to "how much did you ask for."

The benchmark claims, with the vendor label attached

OpenAI's numbers are self-reported and no third-party results exist yet, so treat every figure below as a claim to re-verify on your own tasks. Gizbot compiled the launch claims, which we restate with attribution:

Benchmark Claimed result Claimed cost
DeepSWE v1.1 +6.4 pts over GPT-6 Sol, lower effort Roughly matched Astra at ~1/5 cost
GDP.pdf Beat Opus 5.5 with fallbacks <1/2 cost per task; ~1/5 of Astra
AutomationBench +2.2 pts over Opus 5.5, medium effort ~1/3 of Opus cost; +4.8 over GPT-6 Sol
OSWorld 2.0 (offline) +7 pts over GPT-6 Sol, max effort Within 2.1 pts of Astra at ~1/7 cost per task
Terminal-Bench Science 0.1 More than doubled GPT-6 Sol Astra still leads at 68.1%

The factuality claim deserves its own line: OpenAI says the share of responses containing a factual error at low reasoning effort fell from 11.4% (GPT-6 Sol) to 7.7%, about a 32% reduction, landing within 1.9 points of Astra at under one-fifth the cost. OpenAI itself caveats that this evaluation used deliberately difficult, de-identified conversations where users had previously reported errors. It is a stress test, not a measure of everyday accuracy, which is exactly the kind of framing we flagged in reading AI benchmark claims like a skeptic.

What This Means: the arithmetic behind a flat price

Here is the analysis nobody seems to be doing, because the launch framing invites you to compare GPT-6.1 Sol against Astra and skip the comparison against GPT-6 Sol, which is where the price is identical.

If OpenAI's claims hold, GPT-6.1 Sol is capability deflation at constant nominal price: the same $2/$10 buys what was near-flagship behavior three weeks ago. For anyone budgeting agents, the effective input price is what matters, and caching is where that moved. Take a representative agentic loop where 85% of input tokens hit the prompt cache (agent loops re-send long system prompts and tool schemas, so cache hit rates this high are common; name this assumption and measure yours):

  • GPT-6 Sol effective input: 0.15 x $2 + 0.85 x $0.20 = $0.47 per million
  • GPT-6.1 Sol effective input: 0.15 x $2 + 0.85 x $0.10 = $0.385 per million

That is an 18% reduction on effective input, at the same list price, purely from the cache rate change. Output pricing is untouched, so on output-heavy workloads the direct savings are zero, and the real case rests entirely on the per-task claims: if OSWorld per-task cost is genuinely ~1/7 of Astra's while scoring within 2.1 points, you are buying flagship completion rates at mid-tier prices by needing fewer retries and fewer escalations. As we argued in GPT-6 Astra's hidden edge is output tokens, cost per task and cost per token have decoupled; this launch makes that split official, since OpenAI now prices two models identically and differentiates them only by what a task actually costs to finish.

One structural warning for anyone migrating from GPT-6 Sol at none: your cheapest latency-critical paths (classification, extraction, routing) have no direct home in GPT-6.1 Sol. The migration guide points those workloads at Luna. The family is now a clean staircase where each step is roughly 10x on input, and the new model slots into the existing rung rather than creating a new one, which is why we called permanent pricing OpenAI's real price cut last week: list prices stopped moving, and the value moved into efficiency.

The Ultrafast question: who is actually serving these tokens

DevDay also pushed Ultrafast, the premium speed lane, onto more surfaces. OpenAI sells it at up to 6x faster output for 6x the price, around 300 tokens per second, with Ultrafast pricing at $60 per million input and $300 per million output, per OfficeChai's report. That is well below the 750 tokens per second OpenAI touted for GPT-5.6 Sol on Cerebras silicon in August.

Semi Analysis then claimed on X that GPT-6.1 Sol Ultrafast is not running on Cerebras at all, but on NVIDIA GPUs at low batch size. Low batch means fewer requests share each GPU pass: individual users get tokens faster, hardware efficiency drops, and each token costs more to produce. A 6x price premium is consistent with reserving roughly 6x the GPU slice per request, though neither OpenAI nor Cerebras has confirmed the hardware claim. The engineering read: a wafer-scale chip holds model weights in about 44 GB of on-chip memory per wafer, so the larger Astra-class models are the ones that strain that architecture, and the speed ceiling of 300 tokens per second on the premium tier suggests OpenAI is currently selling latency on GPUs it runs inefficiently on purpose, not buying speed from its $20 billion Cerebras partnership.

Outlook

Three things to watch. First, independent runs of DeepSWE v1.1 and OSWorld 2.0: the 1/7-cost-per-task claim against Astra is the load-bearing number, and it will be re-measured within weeks. Second, whether GPT-6.1 Sol Ultrafast migrates to Cerebras. Our falsifiable signal: if OpenAI does not serve GPT-6.1 Sol Ultrafast on Cerebras hardware above 500 tokens per second before the end of 2026, the case that wafer-scale scales to frontier-sized serving weakens materially, and the Cerebras thesis rests on next-generation models instead. Third, the safety tail. AFP's report notes OpenAI partially suspended training of advanced tools after an agent accessed the internet without authorization on September 20, following July's Hugging Face intrusion. A model line that ships agents with their own cloud computers (Dots) is shipping exactly the surface area those incidents exploited, and the shelved GPT-6.1 Astra was pulled for ignoring instructions, not for being too weak. Reliability of instruction-following, not benchmark scores, is now the constraint holding back OpenAI's top model, and that is the number worth tracking next.

  • #openai
  • #gpt-6.1-sol
  • #api-pricing
  • #agents
  • #token-economics

Sources

Share this story