memujo
AI8 min read

From $30K Wafers to Free Tokens: The AI Cost Chain

AI prices fall and AI spending rises at the same time. The chain from $30,000 wafers to free output tokens explains why every saving gets spent.

By Alice

In this article
  1. 01Layer one: the wafer is getting more expensive, on purpose
  2. 02Layer two: memory is where the chain actually bites
  3. 03Layer three: the hardware arms race is a fight over the denominator
  4. 04Layer four: at the API, the savings land as free
  5. 05The chain has a name, and it is not a paradox
  6. 06Our Read
  7. 07Outlook
  8. 08References

Here is a paradox you can observe without leaving your desk. The price of an AI token has fallen by orders of magnitude in three years. The price of everything underneath it, wafers, memory, servers, has risen. And total AI spending is climbing faster than either curve. That combination looks like a contradiction, but it is a single system seen end to end, and once you trace one dollar through it, most of the confusing news of the last year snaps into a coherent picture.

This piece synthesizes the cost chain we have covered piecewise over recent weeks, from TSMC's 2nm wafer pricing to Meta's Arke inference chip to TypeSafe's free-output pricing model. Every number below repeats a figure we previously traced to a primary source, restated here with the same caveats. Nothing is new information. The point is the chain.

Layer one: the wafer is getting more expensive, on purpose

The chain starts at the fab. TSMC's N2 node entered volume production in Q4 2025, and supply-chain reporting has converged on a sticker of roughly $30,000 per 300mm wafer, about 50 percent above 3nm. TSMC's own technology page confirms the milestone and the architecture shift behind it: N2 is the company's first gate-all-around nanosheet node, and the follow-on A16 promises 15 to 20 percent power reduction and up to 1.10x chip density versus N2P.

The counterintuitive part, which our 2nm cost analysis laid out in detail: the wafer price is not the unit price. What a customer buys is a good die, and die cost is wafer price divided by (dies per wafer times yield). Advanced nodes shrink each die, which packs more dies onto the same $30,000 wafer, and early-yield drag erodes that gain. The economics work when density outruns yield loss, and the entire foundry business is a bet that it does. Node transitions survive their own price hikes because each transition sells more capability per dollar, not less. Hold that sentence; it is the whole article in miniature.

Layer two: memory is where the chain actually bites

One layer up, the same pattern is louder. Memory chips crossed 50 percent of global semiconductor revenue this year, a share nobody expected a decade ago, and the reason is structural: transformers are memory-bound, not compute-bound. Our HBM4 cost-curve analysis walked through why each HBM generation consumes roughly multiples of wafer capacity per bit delivered, stacking dies, interposers, and packaging steps onto the same constrained wafer supply that logic dies need.

When memory demand outgrows supply, it does not politely raise memory prices. It reprices the whole system. Nvidia notified its largest customers in August that Grace Blackwell 300 and Vera Rubin 200 server prices would rise 15 to 17 percent for 2027 shipments, per reporting from Bloomberg, Reuters, and The Information collected in our server price piece. Microsoft, Google, and AWS absorb it. The memory shortage is not a component story; it is the tax collector on every token generated downstream.

Layer three: the hardware arms race is a fight over the denominator

Now the interesting part. Every serious player just announced hardware whose entire pitch is dividing the cost of a token by a larger number.

OpenAI's Jalapeño, built with Broadcom, posted at Hot Chips 2026: about 1.9x better tokens per kilowatt than Nvidia's GB200 at 700 watts versus 1,200, and up to 3.6x lower end-to-end latency against GB300, verified in person by SemiAnalysis's InferenceX suite, as we covered in our benchmark breakdown. Nvidia answered with Groq 3 LPX, a 3,400-tokens-per-second interactive accelerator built on the acquired Groq LPU architecture. Meta's MTIA 450 Arke targets first-half 2027 deployment with, per Meta's own numbers, 4.5x bandwidth improvements; Meta's VP Yee Jiun Song gave the blunt reason: when you build gigawatts of capacity, a 30 percent cost increase becomes unacceptable. And MLPerf's v6.1 round put Vera Rubin NVL72 at 3.7x Qwen3-VL throughput over its predecessor with near-linear rack scaling.

Notice the honest caveat we attached to each: Jalapeño's numbers are OpenAI's with one independent verifier; Arke's are Meta's alone; Nvidia's MLPerf entries are self-submitted. Vendor efficiency claims are marketing until a third party reproduces them. But the direction is corroborated by the sheer number of independent parties claiming it. Efficiency is doubling roughly every year at the silicon layer, and nobody can opt out, because a provider that skips a generation pays the 30 percent Meta refused to pay.

Layer four: at the API, the savings land as free

Now follow the dollar to the invoice. GPT-6 Astra launched at a pricing shape our token-economics piece examined: $10 per million input tokens, $50 per million output. Output at 5x input is the old autoregressive tax: every emitted token costs a full decode pass, so the answer is dearer than the question.

The frontier of pricing is now attacking exactly that tax. Google's Gemini 3.8 Live split voice billing into $0.005 per input minute and $0.018 per output minute, versus OpenAI's flat $0.05 voice-layer minute, making a balanced call hour $0.69 against $3.00. TypeSafe's Jev deleted output billing entirely at $0.042 per million input tokens, which we showed is arithmetic, not charity: remove autoregressive decode and the output token stops being a product. And the local-versus-API break-even math we published in September shows the API floor moving fast enough that hardware owners' break-evens keep sliding further out.

Read layers one through four as a pipeline: wafer prices up 50 percent, server prices up 17 percent, memory now half the industry, and simultaneously tokens per dollar improving by multiples per year, with output tokens trending toward free. Every layer's price rises, and every layer's unit capability falls faster. Both things are true at once.

The chain has a name, and it is not a paradox

The 19th-century economist William Stanley Jevons watched Britain's steam engines get more coal-efficient and observed coal consumption rise, because efficiency made coal-powered work cheaper, and cheaper work demanded more of it. Every link above is Jevons with a different unit.

More efficient wafers make chips worth putting in more places, which is why TSMC's $30,000 wafer still leaves capacity sold out. More efficient inference (Jalapeño's tokens per kilowatt) does not shrink the data center; OpenAI builds more gigawatts, because the marginal token got cheap enough to deploy everywhere. Cheaper APIs (Gemini's voice math) do not shrink the bill; they move voice from a luxury feature to a default, and call-center volumes nobody priced before appear. And free output decisions (Jev) do not signal collapse in the price of judgment; they create a category, "smart if-statements," priced too low to be declined, exactly as routing decisions at $0.0004 each get spent where a rules engine once sat.

The chain resolves because there are two different goods in it. The wafer, the memory stack, the server rack: those are capacity, and capacity is scarce, so its price rises. The token, the decision, the voice minute: those are prices paid per use of capacity, and competition forces those down. Jevons is what happens when the price of use falls faster than the price of capacity rises: revenue per unit falls, units explode, and the capacity layer captures more dollars every year. Memory crossing half of all chip revenue is not a shortage story. It is the capacity layer collecting the demand that the application layer's discounts created.

Our Read

Three conclusions, stated as our opinion. One, the right mental model for AI economics is two linked markets, not one: a Jevons-deflated services market stacked on a scarcity-appreciating capacity market. Analysts who quote falling token prices as evidence AI is a commodity, or rising wafer prices as evidence it is a bubble, are each reading half the chain. Both curves are the same demand curve seen at different depths. Two, the durable bottleneck is memory, not compute, and the chain says why: efficiency gains at the logic layer (smaller transistors, better tokens per watt) multiply what a single accelerator can do, which multiplies what it wants to hold in HBM. Every silicon victory shifts cost toward the memory layer, which is why the HBM makers and memory manufacturers are the quiet winners of every efficiency announcement in this piece, and why we keep watching memory as the sector to read. Three, a falsifiable claim: the next visible pricing event will be another deletion of a billing dimension (a modality priced at zero, an input tier flattened), not another rate cut, and the capacity stocks will react to efficiency wins as positive demand news. If efficiency announcements start reading as negative for the capacity layer, that is the signal the Jevons loop has finally broken, and this model needs rewriting.

Outlook

Expect the chain's two halves to keep diverging: more wafer and memory price increases into 2027 (the Nvidia server repricing is one round, not the last), and more free or metered-to-nothing billing dimensions at the API layer as vendors compete for consumption volume. The interesting risk is policy or supply, not economics: nothing inside the chain requires memory to stay scarce, but building the fabs to end the shortage takes years, and until then the capacity layer keeps its pricing power.

The honest summary: the wafer and the free token are the same market. Jevons saw it in coal, and every layer of AI, at absurdly different speeds and with different winners, is repeating it. The question for anyone building on AI is not whether the token gets cheaper, it will. It is whether your use case can afford to say no at $0.0004 a decision. At that price, almost nobody can. That is not a bug in the pricing model; it is the entire model.

References

  1. TSMC 2nm technology page, tsmc.com, N2 volume-production status and A16 roadmap claims
  2. OpenAI Broadcom Jalapeño announcement, openai.com, chip specifications and benchmark claims
  3. SemiAnalysis InferenceX verification of Jalapeño benchmarks, semianalysis.com, independent benchmark run
  4. Meta MTIA scale chips technical post, ai.meta.com, Arke specs and deployment claims (vendor-reported)
  5. Google developer announcement of Gemini audio pricing, blog.google, per-minute voice rates
  6. OpenAI GPT-Live-1 API launch post, openai.com, voice-layer pricing
  7. Bloomberg, Reuters, and The Information reporting on Nvidia server price increases, August 2026
  8. TypeSafe Jev announcement, typesafe.ai, pricing and latency claims (early access, vendor-reported)
  • #inference
  • #semiconductors
  • #token-economics
  • #jevons-paradox
  • #ai-pricing

Sources

Share this story