---
title: "OpenAI's Chip Beat Nvidia on Watts. Per Chip It Lost Two of Three."
description: "A breakdown of OpenAI's first Jalapeno benchmarks, why the per watt framing changes what beating Nvidia means, and what teams shipping agents should take from the numbers buried in the appendix."
author: "Jahanzaib Ahmed"
date: 2026-08-26
category: "trends"
readingTime: "17 min read"
tags: ["ai news", "ai-infrastructure", "openai", "ai-agents"]
canonical: https://www.jahanzaib.ai/blog/openai-jalapeno-chip-inference-latency-per-watt
source: https://www.jahanzaib.ai
---
# OpenAI's Chip Beat Nvidia on Watts. Per Chip It Lost Two of Three.

**Key Takeaways**

-   OpenAI published the first Jalapeño benchmarks on August 25, claiming 1.5 to 1.9 times more work per watt and 1.7 to 3.6 times lower end to end latency than Nvidia comparison systems.
-   Those ratios are normalized by rated chip power. OpenAI says so plainly: "we normalized the results using each accelerator's published chip power rating." Jalapeño is rated at 700 watts against 1,200 for the GB200 and 1,400 for the GB300.
-   Divide by chips instead of watts and the picture inverts on two of the three models tested. Jalapeño produces roughly 0.83 times a GB300's throughput on DeepSeek R1 and 0.77 times on Kimi K2.5.
-   SemiAnalysis, who own the benchmark, wrote that "all numbers are provided to us by OpenAI" and that they "did not run the full suite of InferenceX benchmarks nor have we seen AgentX results." Neither wire story mentioned this.
-   AgentX is the suite that measures multi turn, long context serving. The workload that was run is 8k in, 1k out, single turn, which is the least agent shaped test available.
-   You cannot buy this chip. It deploys inside OpenAI in small volumes at the end of 2026 and ramps through 2027, so it reaches you as a change in someone else's price and latency.

OpenAI put numbers on its custom inference chip on Tuesday at Hot Chips, and the wire copy wrote itself. Faster than Nvidia. More efficient than Nvidia. Beats Blackwell.

I read vendor benchmarks backwards, starting at the appendix. This one's appendix is unusually honest and almost nobody quoted it. It holds a sentence that reframes every headline number above it, plus a chart caption that quietly tells you the comparison is measured in a unit most readers will misread.

None of that makes the chip bad. It is a genuinely impressive piece of silicon and the independent analysts who saw it in the lab say so. But if you're building on OpenAI's API and you read "3.6 times lower latency" as a promise about your own AI inference latency next quarter, you've read it wrong in at least four separate ways, and the vendor's own post tells you three of them.

![OpenAI's August 25 2026 announcement page titled Jalapeno's first results show industry-leading speed and efficiency in AI inference, showing a photograph of the Jalapeno package mounted on a teal development board](https://cdn.sanity.io/images/qajb7q5q/production/d345106783beb494a320800a029ebc475b0b92be-2880x1800.png?w=1200&q=75&auto=format&fit=max)

_The headline claim is speed and efficiency together, which existing systems trade against each other. The appendix that supports it sits below the fold and carries the rated wattages every ratio depends on._

## What did OpenAI actually announce about Jalapeño?

OpenAI released the first benchmark results for Jalapeño, its custom inference ASIC built with Broadcom. The chip was measured on InferenceX, a public benchmark from SemiAnalysis, against Nvidia GB200 and GB300 systems across three open weight models. OpenAI reported 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end to end latency.

Richard Ho, who runs hardware at OpenAI, told reporters on a press call that "the bottom line is that the results show a very, very significant performance advance over state of the art." The framing in OpenAI's post is about escaping a tradeoff: existing systems pick either throughput or latency, and Jalapeño claims both. Ho called it the "best of both worlds" in a briefing reported by The Verge, because AI systems typically "have to make a trade-off between the two."

Two of the three models tested come from outside OpenAI entirely. That matters, because the standard objection to a first party chip is that it only runs first party models well, and the results do not support it.

Deployment got compressed in the coverage. Ho estimated Jalapeño would deploy at the end of 2026 "in very small volumes," ramping through 2027, and OpenAI said it will keep buying Nvidia. Ho described the wider compute strategy as including "very good partners," which is the polite way of saying nothing about the purchase orders changes this year.

One conflict worth flagging, since it sits in the coverage itself. TechCrunch says Jalapeño was "First announced last October." SemiAnalysis says "In June, OpenAI unveiled the chip program in partnership with Broadcom," and The Verge independently dates it to June too. OpenAI's post dodges the question with "Since announcing Jalapeño." That makes it two sources against one, and one of those two is the team that went to OpenAI's labs to see the chip.

## What happens to the AI inference latency claims when you divide by chips instead of watts?

Two of the three throughput results invert. The per watt numbers are real, but every one of them is a ratio with rated power in the denominator, and Jalapeño's rated power is roughly half its competition's. Multiply each side back out by its own wattage and Jalapeño produces less raw throughput per chip than a GB300 on both of the larger models.

Here is the sentence that does it, from OpenAI's post, immediately under the headline chart: "To compare the systems consistently, we normalized the results using each accelerator's published chip power rating. Jalapeño is rated at 700 watts, although its measured sustained power remained at or below 550 watts on the workloads tested."

![OpenAI chart of mixed tokens per second per kilowatt showing Jalapeno at 85,448 versus 44,960 for GPT-OSS 120B, 19,641 versus 11,781 for DeepSeek R1 and 18,195 versus 11,862 for Kimi K2.5, with the power normalization paragraph printed directly beneath it](https://cdn.sanity.io/images/qajb7q5q/production/b3d42512ee87751bd58b5917d2d0a8530f79b175-2880x1800.png?w=1200&q=75&auto=format&fit=max)

_The 1.9x, 1.7x and 1.5x labels are all per kilowatt. The paragraph underneath is where OpenAI tells you the denominator is a rating, not a measurement._

Run the arithmetic on the appendix figures. Jalapeño is rated at 700 watts, the GB200 at 1,200 and the GB300 at 1,400, and those ratings appear in OpenAI's own chart captions.

| Model | Compared with | Per kilowatt | Rated watts | Per chip |
| --- | --- | --- | --- | --- |
| GPT-OSS 120B | GB200 | 1.90x | 700 vs 1,200 | 1.11x |
| DeepSeek R1 670B | GB300 | 1.67x | 700 vs 1,400 | 0.83x |
| Kimi K2.5 1T | GB300 | 1.53x | 700 vs 1,400 | 0.77x |

That last column is mine, not OpenAI's. It is the published per kilowatt ratio multiplied by each side's rated wattage, so on GPT-OSS a 1.90x per kilowatt win becomes a 1.11x win per chip, and on Kimi a 1.53x per kilowatt win becomes 0.77x. Strictly it is per package, which is the unit OpenAI's own chart captions use, and a GB200 package is not a single die.

Be careful what that proves. It does not mean OpenAI cooked the numbers. The normalization is stated in the open, the wattages sit in every chart caption, and OpenAI even flags that Jalapeño's measured draw stayed at or below 550 watts, which makes its own results look worse than a measured comparison would. That is the opposite of hiding something.

The finding is narrower and more useful. Per watt and per chip are different questions with different answers here, and the headline everyone wrote answers only the first.

## Why did OpenAI normalize on watts in the first place?

Because power is the binding constraint on serving capacity right now, and it is not close. SemiAnalysis put it directly: OpenAI "is currently limited by datacenter power, not by budget or floorspace." When you cannot get more megawatts, tokens per megawatt is the number that decides how much product you can sell.

This is where OpenAI's choice of metric is defensible and the reader's instinct is the thing that's wrong. Jensen Huang made the same argument at Computex 2026, quoted by SemiAnalysis: "If you have 1 gigawatt of power, then throughput per watt is revenue." Nvidia repeated it at its own Hot Chips talk: "The data center is power limited today." Both vendors agree on the metric. They disagree only about whose chip wins on it.

I wrote about the physical version of this constraint when Amazon committed to [7.65 gigawatts of generation that will never touch the grid](https://www.jahanzaib.ai/blog/ai-data-center-energy-amazon-off-grid-gas). Interconnection queues run for years while hardware refreshes run for months, which is why operators end up building their own power. Perf per watt is the only lever that moves without a utility's permission.

So OpenAI did not pick a flattering unit. It picked the unit that governs its own business. The problem is that "beats Nvidia" reads to a developer as a claim about speed, and per watt is a claim about economics.

## Who actually ran the benchmark?

OpenAI did. SemiAnalysis verified the runs in person but did not run the suite themselves, and they said so in the first of their three caveats. This is the single most load bearing fact in the story and it appears in neither of the two largest news write ups.

Their wording: "Some caveats on this. First, all numbers are provided to us by OpenAI. We verified the InferenceX runs in person in the lab, but we did not run the full suite of InferenceX benchmarks nor have we seen AgentX results."

![SemiAnalysis article paragraph stating that all Jalapeno numbers were provided by OpenAI, that they verified InferenceX runs in person but did not run the full suite, and that AgentX results have not been seen](https://cdn.sanity.io/images/qajb7q5q/production/b7fbc1c9cf31ef13a9a6ededd1e80eb528f9a7bf-2880x1800.png?w=1200&q=75&auto=format&fit=max)

_SemiAnalysis own the InferenceX benchmark and were invited to the lab. Their caveat paragraph is the disclosure that did not make it into the wire coverage._

Read TechCrunch and The Verge and you would reasonably conclude an independent lab measured a chip. What happened is closer to a supervised demo: the vendor ran its own numbers, the benchmark authors watched, and the benchmark authors published their read of what they saw. That is meaningfully better than a press release and meaningfully worse than an independent test, and the distinction is worth one sentence in a news story.

Credit where it belongs, because SemiAnalysis are not shills here. In the same article they call the Blackwell comparison "somewhat incomplete and unfair," argue Jalapeño "is really competing against chips like Rubin that also use HBM4," and note that "Vera Rubin systems are starting to ship to customers right now, while it will still be some time before OpenAI has anything beyond engineering samples of Jalapeño." That is the benchmark owner telling you the comparison is against last generation.

TechCrunch caught the same edge independently, writing that "by the time Jalapeño reaches full deployment, the competition may have advanced significantly." Worth being precise about attribution there: that is Russell Brandom's observation in TechCrunch's voice, not something Ho said.

![TechCrunch article body noting the comparison is against an Nvidia Blackwell system, that competition may have advanced by full deployment, and quoting OpenAI on minimizing data movement and keeping the KV cache local](https://cdn.sanity.io/images/qajb7q5q/production/b27d3159c50dce87298dd76003221ae19fd706c5-2880x1800.png?w=1200&q=75&auto=format&fit=max)

_TechCrunch flagged the generational gap in its fourth paragraph, and also carried the KV cache detail that explains where the latency win comes from._

## Why does the missing AgentX run matter for agents?

Because AgentX is the suite built to measure agent shaped serving, and it was not run. The workload that was run is 8k tokens in, 1k tokens out, single turn. Agents are the opposite of that: many turns, growing context, heavy cache reuse, and a router deciding where each call lands.

SemiAnalysis describe AgentX as "our preferred suite for comparing chip performance due to the datasets' long context and multi-turn characteristics that reflect the cache behavior of realistic production workflows." Then the warning: "Frameworks that perform well on 8k1k may perform worse on AgentX as real production loads stress components like routers, prefix cache mechanisms, cache management, offload infrastructure, etc."

Sit that next to OpenAI's own pitch. OpenAI's post says the design question was "what hardware would we build if its primary job were serving modern and future language models, especially interactive agents," and that the gains matter "especially for agents, which need to complete many steps in sequence, so delays can compound across an entire task." The claim is about agents. The benchmark is about single turn chat.

I keep seeing this gap in agent infrastructure claims and it is rarely deliberate. Single turn benchmarks are cheap to run and easy to compare. Multi turn benchmarks with realistic cache behaviour are expensive, and they are also where systems fall over. In my experience the first thing that breaks when a demo becomes a product is prefix cache hit rate on turn six, not tokens per second on turn one. The same blind spot is why OpenAI's own disclosure that agent monitoring [costs 20 percent of the compute](https://www.jahanzaib.ai/blog/openai-agent-monitoring-20-percent-compute-overhead) landed as a surprise.

The compounding argument does cut in OpenAI's favour if the latency holds up, and it's the strongest version of their case. Take the end to end latency figures straight from the appendix and run them out over a modest twelve step agent loop.

| Model | One request | Comparison system | 12 sequential steps | Time saved |
| --- | --- | --- | --- | --- |
| GPT-OSS 120B | 1.03 s | 1.80 s | 12.4 s vs 21.6 s | 9.2 s |
| DeepSeek R1 670B | 1.65 s | 5.99 s | 19.8 s vs 71.9 s | 52.1 s |
| Kimi K2.5 1T | 1.56 s | 5.31 s | 18.7 s vs 63.7 s | 45.0 s |

Twelve steps is my number, not OpenAI's, and the multiplication is the naive one. Real loops overlap calls, cache aggressively and spend time in tools rather than tokens. But the shape holds: on DeepSeek R1 the difference between a user waiting twenty seconds and a user waiting seventy two is the difference between a feature and an abandoned tab. That is why the latency column matters more than the throughput column for anyone shipping [multi agent systems](https://www.jahanzaib.ai/blog/multi-agent-ai-failure-modes-anthropic-research).

## What is the 104x number really measuring?

It measures how badly a GB300 degrades at its own fastest setting, not how fast Jalapeño is. The appendix reports "more throughput at previous TBT" of 104.3x on DeepSeek R1, and that comparison is taken at the point where the comparison system has already collapsed to 118 mixed tokens per second per kilowatt. Any number divided by 118 is going to be large.

![OpenAI appendix for DeepSeek R1 showing package TDP of Jalapeno 700 W and GB300 1,400 W in the chart caption, a bar chart where the comparison system falls from 11,781 to 118 mixed tokens per second per kilowatt, and summary cards reading 1.7x, 3.6x, 4.1x and 104.3x](https://cdn.sanity.io/images/qajb7q5q/production/4de08886ef707615b2ec1fbd0fae7b087f2fd544-2880x1800.png?w=1200&q=75&auto=format&fit=max)

_Two things live in this one frame: the rated wattages in the caption, and the blue bars falling from 11,781 to 118 as interactivity rises. The 104.3x is measured against that last bar._

Look at the blue bars in that chart. At peak efficiency the GB300 does 11,781. Push it to 169 tokens per second per user and it does 118, a drop of about ninety nine percent. Jalapeño's equivalent curve falls from roughly 19,641 to 12,258 over the same range, which is the actual finding: it holds its efficiency as you push interactivity, and the competition does not.

That is a good result stated plainly. Stated as 104.3x it invites a reader to think the chip is a hundred times faster, which it is not by any measure in the same document. The same corner case produces 53.7x on GPT-OSS and 56.1x on Kimi.

I was wrong about this on first read, for the record. I saw 104.3x, assumed a typo, and only found the mechanism after pulling the underlying bar values out of the chart. It is a real number that means something much smaller than it looks.

## What is genuinely impressive here?

The efficiency curve, the development timeline, and the fact that a first generation part is competitive at all. SemiAnalysis, who spend most of their year finding holes in vendor claims, wrote that Jalapeño beats every Nvidia, AMD and Google chip they have tested on perf per watt across the models available to them.

Three details deserve more attention than they got. The results were achieved with single token prediction, no speculative decoding and no prefill decode disaggregation, while the Vera Rubin figures they compare against do use speculative decoding, which SemiAnalysis note "leads to a ~3-5x reduction in cost per token." Jalapeño is fighting with one hand down.

Second, the silicon that produced these numbers is already old. SemiAnalysis report a B0 stepping in the fab with "roughly a 25% perf-per-watt improvement over the earlier A0 silicon." Every published figure is A0.

Third, the AI assisted design claim is unusually concrete. For selected GPT-OSS attention and mixture of experts blocks, OpenAI says "AI-generated implementations ran 1.5 to 1.8 times faster than the existing human-expert-written implementations," then immediately caps the claim: "Those figures apply to the selected blocks, not the full model." TechCrunch notes OpenAI's models assisted development, but neither wire story carried the kernel result, and I think it's the more durable story. A chip that is a tractable target for a model to program is a different kind of asset than a chip that is merely fast.

The timelines need care, because two clocks are running. OpenAI says it went "from initial design to tapeout in nine months." SemiAnalysis date the program from the middle of 2024 and put "initial team hiring to manufacturing tape-out in ~16 months," with tapeout in November 2025. Those measure different spans rather than contradicting each other, and the compute bill behind either is the kind of number that [wrecked a $45 billion fund that had the thesis right](https://www.jahanzaib.ai/blog/sec-probe-situational-awareness-ai-infrastructure-costs).

## When does any of this reach people building on OpenAI?

Not this year in any way you'd notice. Deployment starts inside OpenAI's own infrastructure at the end of 2026 in what Ho called "very small volumes," ramping through 2027, and OpenAI has not said how many chips. There is no version of this where you buy one.

That is the structural point. Your AI inference latency and your inference price are line items on someone else's hardware roadmap. When Google's Flash pricing [doubles on January 1](https://www.jahanzaib.ai/blog/gemini-3-7-flash-pricing-doubles-january-2027), that is the same mechanism pointing the other way. When Stripe pays seven billion dollars for [a default routing setting](https://www.jahanzaib.ai/blog/stripe-openrouter-acquisition-llm-routing), it is buying influence over this exact layer.

The practical read: nothing here changes your architecture this quarter. What it changes is the cost curve you should plan against. Serving efficiency is improving fast enough that per token pricing has room to fall, and it is improving specifically at the interactivity end of the curve, which is the end agents live on.

It also raises the cost of being locked to one provider's silicon roadmap. If OpenAI's serving economics improve 1.5x per watt and a rival's do not, that shows up in your bill within a year or two.

## What would I want to see before believing the agent claim?

Four things, and all of them are cheap for OpenAI to publish. An AgentX run, since that is the suite that stresses the components agents actually stress. Results on a current frontier model rather than three open weight models. Measured power alongside rated power, since the post already admits real draw was under 550 watts. And a per chip column next to the per watt column.

That last one is not a gotcha. Both numbers are legitimate and they answer different questions. An operator with a fixed megawatt budget wants per watt. A developer asking whether responses get faster wants per chip and per user. Publishing one and letting the press round it to the other is where the confusion enters.

I've deployed enough agent systems to know which number I'd check first, and it is neither of those. It's tokens per second per user once real concurrency arrives, and the figures SemiAnalysis reported sit at the easy end of that range: over 700 on DeepSeek R1 at concurrency one, roughly 1,400 on GPT-OSS. Those decide whether a user watches a response stream or a spinner, and they're the figures most likely to move once real traffic and real cache pressure arrive.

If you're weighing how much of your stack to build on a single model provider right now, the [AI readiness assessment](https://www.jahanzaib.ai/ai-readiness) walks through the dependency questions this kind of announcement should prompt, including where a provider swap would actually hurt.

## Frequently asked questions

### Did OpenAI's Jalapeño chip beat Nvidia?

On performance per watt, yes, on the three open weight models tested against GB200 and GB300 systems. Per chip the picture is mixed: Jalapeño leads by about 1.11x on GPT-OSS 120B but trails at roughly 0.83x and 0.77x on DeepSeek R1 and Kimi K2.5, because its rated power is 700 watts against 1,200 and 1,400. SemiAnalysis also argue the fair comparison is against Nvidia's Vera Rubin generation, which is shipping now.

### Who ran the Jalapeño benchmarks?

OpenAI ran them. SemiAnalysis, who publish the InferenceX benchmark, verified the runs in person at OpenAI's lab but wrote that "all numbers are provided to us by OpenAI" and that they did not run the full suite themselves. That is more rigorous than a press release and less rigorous than an independent test.

### What is AgentX and why does it matter?

AgentX is SemiAnalysis's benchmark suite for agentic inference, using long context and multi turn datasets that exercise routers, prefix caches and offload infrastructure. It has not been run on Jalapeño. Since OpenAI's stated design goal was serving interactive agents, the absence of an agent shaped benchmark is the largest open question in the announcement.

### Can I buy or rent a Jalapeño chip?

No. Jalapeño deploys inside OpenAI's own compute infrastructure starting at the end of 2026 in very small volumes, ramping through 2027. It reaches developers only indirectly, as changes to OpenAI's API latency, capacity and pricing.

### Does this mean OpenAI stops buying Nvidia?

No. OpenAI's post says it will "continue to widely deploy accelerators from NVIDIA and other partners for both training and inference workloads," and Richard Ho described the compute strategy as including "very good partners." Jalapeño is additive capacity, not a replacement lineup.

### How much faster would my agent actually get?

Nobody can tell you yet, because the published latency figures come from a single turn 8k in, 1k out workload rather than a multi turn agent trace. As a rough shape, the appendix latencies of 1.65 seconds versus 5.99 seconds on DeepSeek R1 would compound to about 20 seconds versus 72 seconds across twelve sequential calls, before any caching or tool time.

### What is time between tokens and why is it in every chart?

Time between tokens, or TBT, is how long a system takes to emit each successive token, and it is the reciprocal of tokens per second per user. It governs whether streamed output feels smooth or stuttering. OpenAI reports minimum TBT of 0.69 milliseconds on GPT-OSS against 1.87 for the comparison system, which is where the "more responsive agents" claim comes from.

> **Sources:** OpenAI reported 1.5 to 1.9 times more work per watt and 1.7 to 3.6 times lower end to end latency, normalized on rated chip power of 700 W for Jalapeño against 1,200 W and 1,400 W for the GB200 and GB300. [OpenAI, Jalapeño's first results (August 25, 2026)](https://openai.com/index/jalapeno-first-results) · [SemiAnalysis, OpenAI Jalapeño: Better Than Nvidia Blackwell (August 25, 2026)](https://newsletter.semianalysis.com/p/openai-jalapeno-better-than-nvidia) · [TechCrunch, Russell Brandom (August 25, 2026)](https://techcrunch.com/2026/08/25/openais-jalapeno-chip-is-built-for-fast-inference-at-scale-benchmarks-show/) · [The Verge, Emma Roth (August 25, 2026)](https://www.theverge.com/ai-artificial-intelligence/984290/openai-jalapeno-ai-chip-benchmarks) · DeepSeek R1, the model carrying the largest reported latency gap, is documented in [its arXiv paper (2025)](https://arxiv.org/abs/2501.12948). The per chip column, strictly per package, and the twelve step latency projection are this site's own arithmetic on OpenAI's published appendix values.

## Related

- [1,200 Agents Built Their Own Message Board. Almost None Thought to Call a Human.](https://www.jahanzaib.ai/blog/openai-hugging-face-incident-report-ai-agent-oversight)
- [OpenAI Put a Number on Watching Its Own Agents. It Is 20% of the Compute.](https://www.jahanzaib.ai/blog/openai-agent-monitoring-20-percent-compute-overhead)
- [Stripe Is Paying $7 Billion for a Default Setting](https://www.jahanzaib.ai/blog/stripe-openrouter-acquisition-llm-routing)

---

Canonical HTML version: https://www.jahanzaib.ai/blog/openai-jalapeno-chip-inference-latency-per-watt
