
On August 25, 2026, OpenAI published benchmarks for Jalapeño, its first custom silicon chip, at the Hot Chips conference. The next day, Nvidia reported $96.2 billion in quarterly revenue, up 106% year-over-year, the best single quarter in the semiconductor industry’s history.
Both things happened within 24 hours. Both are true simultaneously. Hold them together, because the story only makes sense if you do.
What Jalapeño Is, and Why Energy Is the Constraint That Matters Now
Jalapeño is not a GPU. A Nvidia GPU is a general-purpose chip: it handles gaming, 3D rendering, AI training, AI inference, and a dozen other workloads. Jalapeño is an ASIC, an application-specific integrated circuit, designed to do exactly one thing — take a language model and generate responses as fast as possible while consuming as little electricity as possible. It generates tokens. That is its entire purpose.
The reason this matters is that the primary bottleneck in AI infrastructure right now is not money, not physical space in data centres, and not the availability of researchers. It’s electricity. The companies building AI at scale are limited by how many watts they can secure. OpenAI has been pursuing nuclear power agreements and other non-standard energy supply arrangements precisely because the demand for compute has exceeded what conventional power infrastructure can deliver on the timelines the industry requires.
That makes the key metric obvious. For every watt of electricity you give a chip, how many useful responses does it produce? The more responses per watt, the more revenue you can generate from the same power bill. Jensen Huang said exactly this at Computex 2026: compute is now the billable unit, and power is the constraint on how much compute you can deliver.
When OpenAI arrives with a first-generation chip that beats Nvidia on precisely that metric, the whole industry pays attention.
What SemiAnalysis Actually Found
SemiAnalysis, the most respected independent analysis firm in semiconductors, sent analysts to OpenAI’s labs in person to run benchmarks. This is not a corporate press release. Their conclusions carry weight because they derive from direct observation, and because SemiAnalysis has no commercial reason to flatter OpenAI.
Their findings, tested on the InferenceX benchmark across three models (including two OpenAI had not built, DeepSeek R1 and Kimi K2.5): Jalapeño delivers between 1.5 and 1.9 times more tokens per watt than Nvidia’s Blackwell systems (specifically the GB300 platform), while running at 700 watts of total package power against Blackwell’s 1,400 watts. Half the power draw for a superior result. On latency, the improvement ranges from 1.7 to 3.6 times. In interactive workloads, like those that simulate a genuine conversation with a chatbot, the latency advantage becomes four times greater.
SemiAnalysis is honest about the caveat, and it’s an important one. Comparing Jalapeño to Blackwell today is comparing a first-generation newcomer to last year’s Nvidia platform. The fair comparison is with Vera Rubin, Nvidia’s newest platform, which also uses HBM4 memory and began shipping this month. When you run that comparison directly, the performance gap narrows considerably. In total cost per token, SemiAnalysis found the two platforms roughly at parity.
Here is the actual signal: Jalapeño is a first-generation chip with three months of software optimisation behind it, and it is already trading blows with the most advanced inference platform Nvidia has ever shipped — a platform Nvidia has been refining for years. That is not a gap that points in Nvidia’s favour.




