Original briefings. Zero spin.
Every story is an original briefing written from 110+ sources across the spectrum — sources linked so you can verify it yourself.
OpenAI's Jalapeño Chip Beats Nvidia on Efficiency Benchmarks. Nvidia Posts $96.2 Billion Quarter Anyway

OpenAI showed up at Hot Chips 2026 on Tuesday, August 25, with something rare in the chip business: independently verified benchmark data backing up its marketing claims.
The chip is called Jalapeño. It's an inference-only ASIC OpenAI built with Broadcom, and according to OpenAI's own account, design work started in mid-2024 and reached tapeout roughly 16 months later, an unusually fast cycle for custom silicon. Codex was running on it in early 2026. ChatGPT followed not long after, according to Serve the Home's account of OpenAI's Hot Chips presentation.
SemiAnalysis, the chip research firm that runs the InferenceX benchmark, says it sent its own engineers into OpenAI's labs to run the tests directly rather than taking OpenAI's word for it. Across three open-weight models, GPT-OSS 120B, DeepSeek R1, and Moonshot AI's Kimi K2.5, Jalapeño delivered 1.5x to 1.9x more throughput per kilowatt and 1.7x to 3.6x lower end-to-end latency than Nvidia's GB200 and GB300 systems, per SemiAnalysis and confirmed in reporting by Tom's Hardware. On interactive, conversational workloads, the kind that make up most of ChatGPT's actual traffic, the gap widened to 2.1x to 4.1x faster response times, according to Tech Times.
SemiAnalysis, which is not shy about picking winners, put it bluntly: Jalapeño is "beating every Nvidia, AMD, and Google chip we have been able to test."
The Catch Nobody Should Skip
Before anyone declares Nvidia dead, read the fine print, because there is a lot of it, and it matters.
Jalapeño runs at 700 watts. The GB300 it's being compared against runs at 1,400 watts. That's not a small asterisk. Tom's Hardware reported that when you normalize to all-in utility power instead of package TDP, the comparison narrows considerably, from 1.18 kilowatts for Jalapeño to 2.55 kilowatts for the GB300, and the efficiency lead shrinks further, to roughly 1.5x, when the GB300 is tested using multi-token prediction, a production technique Nvidia deployments commonly use but which wasn't used in most of the headline comparisons.
Jalapeño also wasn't tested against Vera Rubin, Nvidia's next-generation platform slated to power the first gigawatt of Nvidia systems OpenAI has agreed to deploy in the second half of 2026, according to Tom's Hardware. And Jalapeño doesn't train models at all. It's inference-only. Training, the far more compute-intensive workload where flexibility and software ecosystem matter more than raw efficiency, remains entirely Nvidia's turf for now.
CNBC's reporting captured the more measured read from industry analysts. TrendForce's Fion Chiu told CNBC that Jalapeño could reduce OpenAI's reliance on Nvidia for inference specifically, but that "for more compute-intensive workloads, like large-scale model training and frontier AI workloads, we believe Nvidia GPUs will remain important given their broad programmability, performance, software ecosystem, and ability to handle a wide range of workloads." Yole Group's Adrien Sanchez told CNBC the chip is a "threat to Nvidia's inference margins, which is the field growing the most at the moment" while acknowledging Nvidia still owns the "vast majority" of AI compute and retains ecosystem lock-in through its CUDA software platform.
Huang's Answer, And a $96 Billion Number
Nvidia didn't wait long to respond. On its August 26 earnings call for the second quarter of fiscal 2027, the company reported $96.2 billion in quarterly revenue, up 106% year over year. CFO Colette Kress told investors that OpenAI's existing and planned commitments still represent roughly 12 gigawatts of Nvidia compute through 2030, according to Tech Times' account of the call.
When Bank of America Securities analyst Vivek Arya asked CEO Jensen Huang how Nvidia should think about investing in companies building potentially competitive chips, Huang didn't dodge it. "Many of these XPUs are inference-specific chips for one cloud or one service," Huang said. "NVIDIA is a platform, an entire AI factory platform that spans the entire AI lifecycle that you can use in any cloud."
Jalapeño is purpose-built for OpenAI's inference workloads on OpenAI's models, largely. Nvidia's GPUs run training and inference across every cloud provider, with a mature software stack that took over a decade to build. Custom silicon from Google, Amazon, Meta, and now OpenAI chips away at Nvidia's inference margins. It does not replace the platform for frontier training, at least not yet.
Both the benchmark and the earnings call are real, and they describe two different battlegrounds. OpenAI has a legitimately fast, legitimately efficient inference chip backed by third-party verification. Nvidia has a $96.2 billion quarter and a 12-gigawatt commitment from the same company that just benchmarked against it. Whether Jalapeño's efficiency edge survives contact with the next generation of Nvidia's Vera Rubin platform, and whether OpenAI's chip volumes scale beyond its own data centers, are the two questions nobody in this story has answered yet. OpenAI says future Jalapeño generations are already in development. Nvidia's next earnings call will be the place to watch whether hyperscaler silicon shows up as a line item Huang has to explain again.
Sources used for this briefing
This briefing was written by UBH's AI agent — these are the reporting inputs it draws on, linked so you can verify.