memujo
AI8 min read

Gemini 4 Argon Debuts Priced to Undercut Astra

Gemini 4 Argon launches at $2 per million input tokens, half its post-intro rate, gated to cyber defenders first. Inside Google's bet.

By Alice

In this article
  1. 01What Google actually shipped
  2. 02The pricing bet: introductory rates are a switching tool
  3. 03What the benchmark split reveals
  4. 04Why This Matters
  5. 05Outlook

Google announced Gemini 4 Argon on September 30, 2026, and then did something frontier labs almost never do: it told you the price it intends to charge after the launch discount, before the model is broadly available. The announcement post by Koray Kavukcuoglu lists an introductory rate of $2 per million input tokens and $10 per million output tokens, and states plainly in a footnote that "after the introductory period expires, the price of $4 per 1M input tokens and $20 per 1M output tokens will apply." Cached input gets a 95% discount off the input price. The model itself goes first to a closed cohort of trusted cyber defenders through Google's Fairwind Program, with paid API customers and Google AI Ultra subscribers to follow at an unspecified date.

A flagship that costs one fifth of the incumbent flagship's sticker price, shipped behind an access gate with its future list price already published, is not a confused pricing page. It is a set of deliberate bets, and the numbers in the post let you read all three.

What Google actually shipped

The headline capability change is not a benchmark score. It is the output ceiling: Argon's maximum output length goes from 64K tokens to 1 million, a 16x expansion that Google calls industry-leading. That is a different shape of model than the industry spent 2025 optimizing for, where the input window ballooned and output stayed scarce.

The announcement grounds the long-horizon pitch in internal deployments rather than evals alone, all of them Google's own claims so far:

Claim Detail Google's stated outcome
Memory optimization Argon agents analyzed fleet-wide profiling telemetry across Google data centers Over 300 TiB freed once rolled out, estimated 500 TiB to 1 PiB total
Codebase migration C/C++ to Rust, scaling from tens of thousands of lines (re2, libgav1) to 800K+ lines (Fuchsia Zircon kernel) Under automated and manual audit before production
libgav1 decoder Agents replaced 32K lines of SIMD code using profile-guided experiments until the compiler auto-vectorized safe Rust 2.7x faster than the Rust port, identical video output
Quantum subroutines Optimizing qubits and gates tradeoffs in bottlenecked subroutines Beat the published baseline by 40% in minutes

The security story is the reason for the gate. Google says Argon can autonomously find, validate, and patch critical software vulnerabilities, and it will ship to trusted defenders without cyber guardrails so they get full capability. Wiz is already using it through its Scan for Good initiative, and Google claims Argon found a critical vulnerability exposing sensitive personal information in hospital software that previous frontier models missed. On CWE-bench v1, Google reports a 68% score, tied for first.

The pricing bet: introductory rates are a switching tool

Read the pricing next to the competition. Per the comparison table in Basic Tutorials' coverage, OpenAI's GPT-6 Astra costs $10 per million input and $50 per million output, so Argon's introductory $2/$10 is exactly one fifth on both axes. Even at the post-introductory $4/$20, Argon is 2.5x cheaper than Astra. And the $2/$10 intro rate is identical to what OpenAI settled on for GPT-6.1 Sol, the mid-tier model OpenAI shipped a day before Argon to defend the value segment.

That last coincidence is the tell. Google is not undercutting Astra at launch; it is meeting Sol at its own price point, then raising the ceiling later, once. A public, dated-in-spirit footnote about the doubling is unusual, and from a platform-economics view it is the smart version of the move. Agent workloads are sticky in a way chat workloads are not: once a team's orchestration harness, eval suite, and retry logic are tuned against a model, switching costs compound weekly. Google has a model nobody outside a vetted cohort can use yet, so the introductory price is buying a head start in the integration race. Teams that wire Argon into their pipelines at $2/$10 are unlikely to rip it out at $4/$20, because by then the doubling is known, budgeted, and already priced into their per-task math.

The cached-input number is the other lever, and it is aimed at the same target. Cached input at 95% off means $0.10 per million tokens, the same effective cache price GPT-6.1 Sol introduced. Agentic traffic is dominated by repeated prompts: a harness that resends a 200K-token codebase context on every tool call pays full price once and pennies on every re-send.

Labeling the arithmetic below as an illustration with stated assumptions, using published list prices only: a workflow with a 200K-token context resent 10 times, one cold read plus nine cache hits, and 50K output tokens totals roughly:

Model Input cost Output cost Total
Argon intro (200K + 9x200K cached, 50K out) $0.40 + $0.18 $0.50 $1.08
Argon post-intro $0.80 + $0.36 $1.00 $2.16
GPT-6 Astra $2.00 + $0.90 $2.50 $5.40

Assumptions: no long-context surcharge applies, cache hit rate is perfect after the first pass, and Astra's own cache discount is ignored since its cache-read price is not in the sources opened for this piece. The ordering, not the precise totals, is the point: an agent-shaped workload is 2 to 5 times cheaper on Argon across its entire announced price range, before anyone debates a single benchmark.

What the benchmark split reveals

Google claims state of the art on DeepSWE v1.1 (77.9%), AutomationBench (51.3%, first place), LVBench (91.7%), and the top of the Vals Index, which weights finance, coding, legal, and tax work by US GDP contribution. Cross-model numbers reported by Basic Tutorials put Argon ahead of GPT-6 Astra on GraphWalks (84.2% vs 71.8%), Vals Finance (65.4% vs 53.5%), and LVBench (91.7% vs 87.5%), and tied with Astra on CWE-bench v1 at 68%.

Here is the pattern a data scientist should pull on: Argon does not win everywhere. The same reporting shows Astra ahead on Terminal-Bench Science (68.1% vs 57.6%) and FrontierSWE v2 (65.5% vs 55.0%). Score the wins and losses by what each benchmark stresses, and the split is almost too clean. Argon's largest margins land where sustained trajectory length is the bottleneck: GraphWalks, long-horizon context navigation, is a 12.4 point gap, and AutomationBench measures end-to-end multi-step execution. The losses land where the task is a hard, short-horizon reasoning spike: science terminals, frontier software puzzles. That is exactly what you would predict from the single spec Google emphasized. A 16x output expansion is an architectural commitment to long trajectories, so the model should dominate tests of stamina and be merely good at tests of raw single-shot difficulty. Treat this as an inference from the score table, not a confirmed mechanism: no architecture details or third-party evals of Argon exist yet, and the caveat I have made on benchmark tables before, in Reading AI Benchmark Claims Like a Skeptic, applies doubly when the eval cohort is Google's own testers.

The software-engineer read on the internal claims is similar. The libgav1 result is the most falsifiable thing in the post, and therefore the most credible: 32K lines of SIMD replaced with safe Rust that the compiler vectorizes, 2.7x faster than the prior Rust port, bit-identical output. That is a measurable, public-repo outcome any team can try to reproduce the moment the API opens. The 300 TiB memory win is the least verifiable (fleet telemetry nobody sees), but note the engineering framing: the agents ran rounds of profile-guided experiments and studied compiler output, which is an agentic search loop over code transformations, not a single generation. That loop is why the 1M output ceiling matters; thousands of experiment transcripts in one trajectory is precisely the workload a 64K model cannot hold.

Why This Matters

Three concrete consequences for builders.

First, the gated launch is a template, and it is the security half that will be copied. Releasing a frontier model to vetted defenders without guardrails, while the general public waits, inverts the old rollout order where consumers got the model and enterprises got it later. Google frames it as safety (the post describes activation-monitoring for misuse, chain-of-thought misalignment monitors, and hardened sandboxes per its agent control roadmap, and says it is engaged in the US government's voluntary pre-release access process). Whatever the motive, the effect is that the first frontier model whose headline capability is autonomous vulnerability patching ships to blue teams before anyone else. Given that Argon's predecessor niche was already covered by Gemini 3.8 Flash Cyber, this cements Google's two-track security play: a small fast cyber model that is public, and a frontier one that is not.

Second, the price war has moved to a third axis. OpenAI's token economics on Astra set the flagship sticker; GPT-6.1 Sol matched it at one fifth with flat pricing; Argon now matches Sol's price with a flagship claim and a published future price. The competition is no longer price per token or even cost per task. It is cost per integration: whoever is cheapest to keep during the next 90 days of building wins the workload.

Third, the Astra-vs-Argon benchmark split should recalibrate expectations for the coming week of third-party evals. If independent numbers preserve the shape (Argon winning stamina tests, Astra winning spike tests), the correct 2026 posture is a router, not a flagship: long-horizon migration and audit workloads to Argon, dense scientific reasoning to Astra, and per-task cost measured on your own traces rather than the 2026 release tracker headline scores.

Outlook

The open questions are all in the gate. Google gives no date for general availability, so the introductory clock effectively starts for everyone at the same moment the API opens, and the $4/$20 doubling will land on roadmaps that were built assuming $2/$10 forever. My bet, falsifiable and specific: if Google holds the introductory price for less than 60 days after broad API access begins, the intro rate was purely a switching subsidy and flag it as such; if it holds longer than six months, read that as Google deciding the real constraint on Argon is serving capacity, not safety review, and expect the pricing footnote to quietly disappear.

The other signal to watch is the guardrail question. Google says trusted defenders get Argon without cyber guardrails and that its internal benchmark found vulnerabilities across codebases in 20 programming languages. The first public report of an Argon-class model weaponized for offense, or a confirmed autonomous patch shipped to production by a Fairwind partner, will tell you whether gated frontier release is a durable policy or an expensive press conference. Either outcome resets what every lab's next launch announcement has to contain. For now, Argon is a frontier model you can price but not call, and the most honest summary is the one the numbers support: the strongest stamina claim of the cycle, priced to be adopted before it can be independently checked.

  • #google
  • #gemini-4-argon
  • #api-pricing
  • #cybersecurity
  • #ai-agents

Sources

Share this story