memujo
AI6 min read

OpenAI Jalapeño Chip Beats Nvidia Blackwell at Hot Chips 2026

OpenAI's custom inference ASIC outperforms Nvidia's GB300 by up to 1.9x on throughput per kilowatt, powered by HBM4 and co-designed with Broadcom.

In this article
  1. 01What Jalapeño actually is
  2. 02The benchmarks: what the numbers mean
  3. 03Why this matters: power is the real constraint
  4. 04The system architecture: a rack, not a chip
  5. 05What happens next

OpenAI revealed the first real-world benchmarks for Jalapeño, its self-designed AI inference chip, at Hot Chips 2026 on Tuesday. The results: the 700-watt ASIC beats Nvidia's 1,400-watt GB300 by up to 1.9x on tokens per kilowatt and delivers up to 3.6x lower end-to-end latency, according to the company's own announcement and independent verification from SemiAnalysis.

The announcement marks a turning point in the race to reduce dependence on Nvidia hardware. OpenAI has spent over two years quietly building Jalapeño from scratch with Broadcom, and the chip is now running frontier models including GPT-OSS, DeepSeek R1, and the 1-trillion-parameter Kimi K2.5. SemiAnalysis visited the labs and ran the InferenceX benchmark suite in person to independently verify the numbers.

What Jalapeño actually is

Jalapeño is not a raw accelerator. OpenAI frames it as a complete inference platform, encompassing the chip, the host CPU pair, and the rack-level system. The architecture was conceived in late 2024, RTL was frozen in 2025, and the first silicon (A0 stepping) returned to the lab early in 2026. The current B0 stepping delivers 13.4 PFLOPs of MXFP4 on a single reticle-sized compute die manufactured on TSMC's N3P process.

The chip uses HBM4 memory with an aggregate bandwidth of over one petabyte per second across 128 chips. That is staggering on paper. A bandwidth-only ceiling implies between 1,000 and 2,000 tokens per second per user without speculative decoding. OpenAI acknowledges the real-world system lands well below that theoretical maximum, which is honest and worth noting.

The development cycle from initial design to tapeout took roughly 16 months. That timeline is notable precisely because it relied on AI-accelerated design tools, including an internal version of Codex that wrote and optimized the chip's kernels. OpenAI even ran Doom at 36 FPS on the chip as a proof-of-concept, ported using nothing but Codex prompts.

The benchmarks: what the numbers mean

OpenAI tested Jalapeño against Nvidia's GB200 (1,200W), GB300 (1,400W), and AMD's MI355X (1,400W) using SemiAnalysis's InferenceX benchmark suite, which normalizes results to each accelerator's published package TDP. All comparisons run Jalapeño in single-token prediction mode, while the Nvidia baselines use multi-token prediction, which is a significant advantage for the competition.

On GPT-OSS 120B, Jalapeño delivers about 1.9x higher peak mixed tokens per second per kilowatt and roughly 1.7x lower end-to-end latency compared to the GB200 at matched operating points. At the previous best time between tokens, it achieves approximately 53.7x more throughput.

The DeepSeek R1 670B results are more striking. Against the GB300, Jalapeño delivers about 1.7x higher peak mixed tokens per kilowatt per second and 3.6x lower end-to-end latency. At the previous best time between tokens, it offers roughly 104.3x more throughput. This is still with single-token prediction only, no speculative decoding, and no prefill-decode disaggregation.

The Kimi K2.5 1-trillion-parameter model shows the largest relative gap in the test set. Jalapeño lands at about 1.5x higher peak mixed tokens per second per kilowatt and 3.4x lower end-to-end latency compared to the GB300, with roughly 56.1x higher throughput than the previous best time-between-tokens measurement.

OpenAI also claims sub-millisecond token-to-token latency on frontier models at economical throughputs. See also: OpenAI recently ramped data operations as it scales its infrastructure.

Why this matters: power is the real constraint

OpenAI designed Jalapeño around two metrics: time to last token for user experience, and tokens per joule for inference efficiency. This is the right framing. Data centers are power limited, not budget limited. Grid interconnection delays, cooling capacity, and UPS design create hard ceilings on how quickly an operator can add capacity, regardless of how much money they have.

As Jensen Huang said at Computex 2026: "If you have 1 gigawatt of power, then throughput per watt is revenue." That sentence captures the entire economics of AI infrastructure right now. Every extra token per megawatt is revenue, and every watt saved is a path to deploying more models without waiting on utility upgrades.

Nvidia itself acknowledged this at Hot Chips 2026 while presenting Vera Rubin. The Vera Rubin NVL72 delivers 5.4x the performance per megawatt of the GB200 NVL72. Vera Rubin is already shipping to customers. That means Jalapeño's real competition is not the Blackwell generation that is already in production, but the Rubin platform that is just starting to arrive. SemiAnalysis noted that comparing Jalapeño to Blackwell is somewhat unfair, because a purpose-built ASIC is always going to outperform a general-purpose GPU on its narrow target workload.

The system architecture: a rack, not a chip

Jalapeño is only half the story. The full system consists of a host CPU rack and an ASIC rack, built with Celestica. Each host tray houses two Turin-class AMD EPYC CPUs with 1.5TB of DRAM. Each ASIC tray contains 8 Jalapeño chips, making 128 chips per rack. The total power draw is about 160kW, roughly the same as a double-wide GB300 rack.

The chips communicate via 8 external PCIe DAC cables per tray and a copper backplane with Tomahawk 6 switch ASICs. The local domain connects 128 XPUs within the rack, while the global domain connects up to 16 racks or 2,048 XPUs using a hybrid of copper and optical interconnect. OpenAI aims to scale to 100MW in production, with a longer-term target of 10 gigawatts.

Crucially, OpenAI chose not to use prefill-decode disaggregation on these chips. NVIDIA and AMD benefit significantly from splitting prefill and decode workloads across separate pools, but that approach assumes traffic patterns are stable and predictable. Real production traffic changes throughout the day, and a fixed hardware split means chips sit idle whenever the traffic mix shifts away from their assigned phase. A unified system keeps every device available for whatever request arrives next.

What happens next

OpenAI is partnering with neoclouds for deployment and gathering reliability data with datacenter partners through January 2027 while optimizing rack installation time. A second-generation chip is reportedly approaching tapeout, expected within months, using a TSMC 3nm-class process. Concept work on a third generation is already underway, Bloomberg reported.

The implications for Nvidia are substantial. The company recently agreed to backstop up to $105 billion in financing to support OpenAI's data center expansion, including plans to lease a 10-gigawatt campus in Ohio from SoftBank. Nvidia's own Vera Rubin is competitive, but if OpenAI's in-house silicon proves reliable at scale, it reduces OpenAI's dependence on Nvidia's hardware supply chain and pricing.

Anthropic is also co-designing custom inference chips to bypass Nvidia, and Google could build more AI accelerators than Nvidia sells in 2028, according to one analyst. The trend is clear: the biggest AI labs are building their own silicon. The question is no longer whether they will succeed, but whether they can scale production fast enough to make it matter.

HBM capacity remains the bottleneck. Samsung, SK Hynix, and Micron have sold their HBM supply through 2027, and Nvidia is reportedly testing cut-down Rubin Ultra configurations with as little as 192GB because memory is unavailable. If Jalapeño ships in volume, it will add competition for TSMC's CoWoS packaging capacity and the HBM supply chain that already cannot meet demand.

The broader signal from Hot Chips 2026 is that AI inference is becoming a dedicated domain where custom silicon can win. GPUs were never designed for the specific access patterns of LLM inference. As OpenAI demonstrates with Jalapeño, when you design from first principles for that workload, the efficiency gains are hard to ignore.

  • #openai
  • #jalapeno
  • #asic
  • #nvidia
  • #hot-chips
  • #broadcom

Sources

Share this story