Back to blogTechnology

OpenAI Jalapeño: the chip challenging Nvidia in inference efficiency

OpenAI releases Jalapeño benchmarks: its in-house chip delivers up to 1.9x more work per watt and 3.6x lower latency than Nvidia's Blackwell.

Published onSeptember 01, 20266 min readFabian Martinelli
Share
OpenAI Jalapeño: the chip challenging Nvidia in inference efficiency

The day OpenAI entered the silicon business

On August 25, 2026, OpenAI published on its official blog the first benchmark results for Jalapeño, its inference chip developed in partnership with Broadcom (NASDAQ: AVGO). The following day, August 26, the numbers were presented at Hot Chips, a conference held at Stanford University — a stage historically reserved for major semiconductor manufacturers. The message was unambiguous: OpenAI is no longer just a model company. It now competes at the silicon layer.

This changes the game. Not because Jalapeño makes Nvidia obsolete — it doesn't, at least not yet. But because it demonstrates, for the first time with verifiable numbers, that a hyperscaler can design inference hardware that is competitive with the absolute market leader. For anyone running businesses that depend on AI APIs, this matters directly to the cost-per-token equation.

What Jalapeño is and how it was built

Jalapeño is OpenAI's first custom chip, designed exclusively for LLM inference — that is, for running already-trained models, not for training them. The distinction matters: training remains Nvidia's territory, with its clusters of Blackwell GPUs. Jalapeño targets the other side of the operation: delivering real-time responses to users and agents.

The project began in mid-2024. The partnership with Broadcom was announced in October 2025, and the chip program was formally introduced in June 2026, described by OpenAI itself as built "from the ground up for current and future LLMs across the industry." From the initial team hire to manufacturing tape-out, the process took approximately 16 months — a remarkably aggressive pace for the semiconductor industry. OpenAI itself used its models to assist in the chip's development.

One technical detail worth noting: the published benchmarks were run on the chip's A0 stepping — an engineering sample. The version currently in production is the B0 stepping, which offers approximately 25% more performance per watt than the tested A0. The disclosed numbers are, in practice, already conservative relative to what will be available in production.

Specifications that define the positioning

The Jalapeño B0, in a single reticle-sized die, delivers 13.4 PFLOPs of MXFP4 compute, manufactured on the TSMC N3P process for the compute die and TSMC N3E for the I/O chiplet. The nominal TDP is 700 watts, but tested workloads sustained average power consumption at or below 550 watts.

The memory choice is strategic: the chip uses HBM4, with 216 gigabytes per accelerator and a bandwidth of 15.4 TB/s per package — which, according to SemiAnalysis, surpasses all accelerators currently using HBM3E. OpenAI is among the earliest HBM4 adopters in the market, alongside Nvidia and AMD, and ahead of TPU and Trainium programs.

At rack scale with 128 chips, the system delivers 1.7 exaflops of 4-bit compute and 27.5 terabytes of HBM4 memory. The full pod supports up to 2,048 ASICs in a global multi-rack scale-up domain.

The benchmarks: what the numbers say — and what they don't

The methodology used was the SemiAnalysis InferenceX benchmark, a public framework that measures the complete end-to-end cycle of serving an AI request. SemiAnalysis analysts were physically present at OpenAI's labs to observe the measurements — which lends directional credibility to the results, even though the numbers are self-reported by OpenAI.

The three models tested, with their respective Nvidia comparison systems:

  • GPT-OSS 120B (vs. Nvidia GB200): throughput per kW ~1.9x higher; latency reduced from 1.80 s to 1.03 s — a 1.7x improvement. The model reached approximately 1,400 tokens per second per user.
  • DeepSeek R1 670B (vs. Nvidia GB300): throughput per kW ~1.7x higher; latency dropped from 5.99 s to 1.65 s — a 3.6x improvement. The chip exceeded 700 tokens per second per user at concurrency 1.
  • Kimi K2.5 1T (vs. Nvidia GB300): throughput per watt ~1.5x higher; latency from 5.31 s to 1.56 s — a 3.4x improvement.

For highly interactive workloads — the profile of ChatGPT and AI agents — the advantage was even greater: 2.1x to 4.1x performance gains.

The caveat that cannot be ignored

SemiAnalysis itself flagged that the comparison is partially incomplete and unfair: Jalapeño uses HBM4, while Nvidia's GB200 and GB300 systems operate on HBM3E — a previous-generation memory. The more balanced comparison would be against the Nvidia Vera Rubin platform, which is not yet widely available. Rubin has a TDP of 900 W to 1,150 W per die and delivers 17.5 PFLOPs in dense NVFP4 — versus the 13.4 PFLOPs of Jalapeño B0 in MXFP4. On cost per token, an analysis by financefeeds.com points to near parity between the two, with Jalapeño still "slightly ahead" in performance per megawatt — but with the caveat that Rubin already incorporates speculative decoding optimizations that Jalapeño does not yet have. (The original source for this specific passage contains partially imprecise excerpts; the data should be read with caution.)

Adrien Sanchez, technology analyst at Yole Group, was direct in comments to CNBC: "A chip designed by a hyperscaler can now match or exceed Nvidia's Blackwell GPUs in inference efficiency."

What changes in practice — and when

Before any operational enthusiasm: Jalapeño is not commercially available. As of September 1, 2026, OpenAI holds only engineering samples. Internal deployment in small volumes is expected by late 2026, as stated by Richard Ho, OpenAI's head of hardware, who explicitly described "very small volumes" for that period. Full production ramp-up will occur gradually throughout 2027, with the bulk of capacity expected by the end of that year.

On the platform's future: Richard Ho confirmed that the second-generation chip is "deep into development." Independently, OpenAI and BigGo Finance reported that a third generation is already underway. OpenAI describes Jalapeño as a multi-generational platform — chips, models, and memory developed together — making clear that this is a long-term investment, not a one-off project.

Why this matters to business decision-makers

We are facing a structural inflection point: major AI labs are verticalizing all the way down to the silicon layer. This compresses the differentiation space for traditional hardware vendors and creates real pressure on the margins and road maps of established players. For SMBs adopting AI, the practical signal is that inference costs will continue to fall, and the most capable models tend to become progressively more accessible as the operating cost structure of major labs improves.

The strategic question is not whether you will use AI in your operations. It is whether you will be positioned to take advantage of the next wave of cost reduction — or whether you will arrive late, as many did to the previous one.