OpenAI shipped GPT-6 Astra on Thursday with a benchmark table that looks competitive on paper, but the real story hides in the footnotes where it reports API cost. For the first time on a flagship release, OpenAI is quantifying the efficiency tradeoff against every competitor in the same row, and the headline number is not accuracy. It is output tokens. On the professional workloads that actually drive agent billables, Astra uses up to 65% fewer output tokens than Anthropic's Claude Fable 5.1 and roughly a third fewer than its own predecessor, GPT-5.6 Sol.
This is a different article than the one written earlier this week about Astra crossing the Critical cybersecurity threshold. That post covered the safety gating. This one covers the economics of a model that has apparently been trained to stop writing before it finishes thinking out loud, and why that behavior matters more to an enterprise buyer than a one-point bump on a research benchmark.
What the release actually claims
GPT-6 Astra is rolling out to a limited set of organizations first, then expanding to ChatGPT Plus, Pro, Business, and Enterprise users over the coming days, alongside the API, Microsoft Azure, and AWS Bedrock. The API pricing is standard at $10 per million input tokens and $50 per million output tokens, with a Fast mode that doubles both speed and price. The standard input and output rates are not dramatic, but the output rate is where the token math bites, because agents that do real work spend most of their budget generating text, reasoning traces, and tool calls rather than reading context.
The benchmark suite is unusually broad for a launch, spanning computer use, professional document work, software engineering, science, and cybersecurity. OpenAI is treating this release as a step change rather than an incremental bump, and the internal numbers support that framing for most categories. The model saturates FrontierMath Tier 4 at 98% and ARC-AGI-3 at 99.9%, and it reached human parity on the action-efficiency levels of ARC-AGI-3. On Terminal-Bench Science, it scores 64.6%, well ahead of Claude Fable 5.1 at 52.6%. On the Agents' Last Exam, a measure of complex real-world professional tasks, it reaches 59.3% against Claude Opus 5 at 55.5% and GPT-5.6 Sol at 53.6%.
Those are strong results, and they are the ones journalists will lead with. But accuracy parity is the commodity now. Every model in the comparison table clears the same difficulty threshold on academic and coding evals. The differentiator has shifted to efficiency, and OpenAI's own cost columns are the most honest part of the release.
The token economics that actually matter
Here is the data that matters for anyone pricing a production agent. OpenAI reports estimated API cost per task across several benchmarks, and the savings compound aggressively on output-heavy workloads.
| Benchmark | GPT-6 Astra | GPT-5.6 Sol | Claude Fable 5.1 | Claude Opus 5 |
|---|---|---|---|---|
| Terminal-Bench Science 0.1 | 64.6% | 22.4% | 52.6% | 30.0% |
| Agents' Last Exam | 59.3% | 53.6% | - | 55.5% |
| Terminal-Bench 4.0 | 57.9% | 37.3% | 55.8% | 19.1% |
| BenchCAD | 95.9% | 83.3% | 84.3% | - |
| GPQA Diamond | 96.0% | 94.6% | 93.7% | 93.7% |
| ARC-AGI-3 | 99.9% | 97.8% | 30.2% | - |
The last row is the anomaly worth explaining. Astra scores 99.9% on ARC-AGI-3 while Claude's reported number sits at 30.2%. OpenAI notes this reflects a different harness, and the two organizations grade the benchmark differently, so the headline gap is not a fair apples-to-apples comparison. Treat the ARC-AGI-3 delta as a measurement artifact, not a capability cliff.
The real economics live in the professional and science workloads, where Astra repeatedly hits higher scores at lower estimated cost. On Agents' Last Exam, OpenAI states Astra uses roughly 65% fewer output tokens than Opus 5 at the highest scoring settings. On BenchCAD, Astra's estimated API cost is about 86% lower than Fable 5.1. Those percentages are not small. They are the difference between an agent that pays for itself and one that requires a headcount to run alongside it.
The mechanism is visible in a feature OpenAI introduced alongside the model: context notes. Historically, long-running agents use compaction to summarize work when the context window fills, and every compaction throws away information about why a fix failed or how a component behaves. Astra can keep notes across context windows, and earlier windows remain searchable. From an engineering standpoint, this reduces the token tax of context management. An agent no longer has to regenerate a summary it already wrote, and it no longer re-reads the entire conversation history to find a requirement. That is a direct reduction in cumulative output tokens across a session, which is the line item that scales with agent usage.
A data-science read on the efficiency move
The token savings are not accidental. They are the visible symptom of a training objective that has shifted toward action-efficiency rather than verbose reasoning. OpenAI's own framing on ARC-AGI-3 uses the phrase "how efficiently it learns to do so," and the benchmark itself measures the ratio of actions to task completion. A model that solves a problem in four tool calls instead of twelve is doing the same work with a third of the output budget.
From a data-science perspective, this mirrors a broader trend in reinforcement learning for agents. The field moved away from reward signals that graded only the final answer toward signals that also rewarded the path. A model trained purely on outcome will happily generate long reasoning chains, because more computation never hurts the score on a held-out test. A model trained on efficiency has a hard incentive to compress. The cost columns in this release are the economic version of that same pressure, because every wasted token is revenue left on the table at a $50-per-million output rate.
There is a measurement caveat worth stating plainly. OpenAI reports the maximum score at any effort level, and it does not publish the token cost at every effort setting. A model can look efficient at its best configuration while consuming excessive tokens at the configuration an enterprise actually chooses. The cost figures are estimates, and OpenAI explicitly notes that production ChatGPT output can differ from the API due to system prompt and tool differences. Treat the percentage savings as directional, not contractual.
An engineering read on what this means for deployment
For teams building agents, the efficiency shift changes three operational decisions.
First, the output token rate is now the dominant cost driver. At $50 per million output tokens, a single agent session that generates two million output tokens costs $100. A session generating six million costs $300, and the third one is pure margin erosion. A model that halves output volume for the same task does not just cut cost, it cuts latency and raises throughput in the same motion, because the model finishes sooner and the downstream system can serve more requests per second.
Second, context notes change how you design long-horizon agents. The ability to preserve and search earlier context without full compaction means an agent can accumulate state across a multi-hour session without the token cost growing linearly with conversation length. That is a structural change to the economics of agentic work, not a marginal improvement. Systems that previously required human babysitting because the context window kept collapsing become economically viable at scale.
Third, the efficiency move raises the stakes on effort-level configuration. OpenAI lets callers choose low, medium, and high effort, and it has confirmed that higher effort buys more verification and more code-execution iterations. The question for an engineering team is no longer whether to use the model, but which effort level per task class. A cheap, fast effort level for triage and a high effort level for final verification is the pattern that preserves the token economics while keeping quality.
Why This Matters
Astra is the clearest signal yet that the frontier has moved past raw capability and into efficiency. Every model in the comparison clears the same eval bar, so accuracy is table stakes. The battle is now over tokens per completed task, and that is a battle enterprises can measure on their invoice.
The practical implication is that the right model is no longer the one with the highest benchmark score. It is the one that completes your specific agent workflows at the lowest output-token cost, and that varies by task. Astra leads on professional document work, computer use, and science workflows, but Anthropic's models remain competitive on certain coding and browser-use benchmarks, and Google's Gemini 3.8 Flash appears on several rows. A portfolio approach, routing tasks to the cheapest capable model per category, is the rational response to a market where every frontier model is good enough and efficiency is the differentiator.
Outlook
Expect OpenAI to keep publishing cost columns. The token economics framing will become standard as buyers shift from "what can the model do" to "what does it cost to do it at scale." The context-notes feature may quietly become the default, because it changes the economics of long sessions in a way that users feel on every invoice. And the effort-level configuration will mature into task-class routing, where an agent automatically picks the cheapest setting that meets a quality bar.
For now, the release is a strong model with a genuine efficiency edge, but the savings are vendor-reported and task-specific. Verify them against your own workflows before re-architecting your agent stack around the headline percentages. External sources: OpenAI, the full GPT-6 Astra release post with the complete benchmark table, cost columns, and pricing | OpenAI Deployment Safety Hub, the GPT-6 Astra system card with the safeguarding and cybersecurity evaluation details.
Internal links: how Astra crossed the Critical cybersecurity threshold and what the gating meant and how OpenAI split its custom-chip fabrication to Samsung to reshaping compute economics.