READ. SCROLL. LISTEN.

Original briefings. Zero spin.

Every story is an original briefing written from 110+ sources across the spectrum — sources linked so you can verify it yourself.

← Back to headlines

OpenAI Says Its Jalapeño Chip Beats Nvidia on AI Response Speed

OpenAI Says Its Jalapeño Chip Beats Nvidia on AI Response Speed
OpenAI unveiled benchmark results Tuesday claiming its custom Jalapeño chip beats Nvidia's Blackwell systems on speed and power efficiency for AI inference. The chip is not shipping in any real volume until late 2026 at the earliest, and OpenAI itself says Nvidia remains a core partner, so treat the victory lap with appropriate skepticism.

OpenAI showed off benchmark numbers Tuesday for Jalapeño, its first custom-built chip, at the Hot Chips conference. The company says the chip beats Nvidia's current top-end systems on the metrics that matter most for running AI models day-to-day: speed and power draw.

According to OpenAI's own blog post, Jalapeño delivered 1.5 to 1.9 times more AI work per watt and 1.7 to 3.6 times lower end-to-end latency than the comparison systems, tested across three models: GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 1T. For highly interactive workloads, OpenAI claims a 2.1 to 4.1 times performance gain.

Richard Ho, OpenAI's head of hardware, told reporters on a press call that AI systems typically force a tradeoff between speed and volume. "Jalapeño can serve more AI work per unit of power, while also returning responses more quickly," Ho said, according to TechCrunch. "It's very efficient to serve a lot of customers, but it can also be very low latency."

Who actually ran the numbers

This wasn't just OpenAI grading its own homework. The company let SemiAnalysis, an independent chip-analysis outlet, into its labs to run the InferenceX benchmark suite itself. SemiAnalysis reported that Jalapeño beat every Nvidia, AMD, and Google chip it has tested on multiple open-source models, calling the results an "immediate contender" against flagship GPUs given the chip's use of HBM4 memory.

SemiAnalysis also pushed back on a common media framing. "A lot of the media coverage of this chip has followed a few throwaway comments from OpenAI that claim the chip will be optimized for their models in a way that other chips are not," the outlet wrote. "This is wrong. Jalapeño is a generalized inference chip capable of running all sorts of models." As a demonstration, OpenAI reportedly showed SemiAnalysis engineers running Doom on the chip, ported over using AI coding tool Codex.

The comparison point matters here. The benchmark measured Jalapeño against Nvidia's GB200 and GB300 Blackwell systems, the best publicly available results at the time. Nvidia's next-generation Rubin chips are on the way, and TechCrunch noted that "by the time Jalapeño reaches full deployment, the competition may have advanced significantly."

Timeline and design partner

Jalapeño is an Application-Specific Integrated Circuit, or ASIC, developed in partnership with Broadcom. OpenAI first announced the program in October 2025 and unveiled more specifics in June 2026. Design work reportedly started in mid-2024, meaning OpenAI went from hiring its chip team to a manufacturing tape-out in roughly 16 months, a fast timeline for a first-generation chip, according to SemiAnalysis.

OpenAI says its own AI models helped design the chip and are now helping engineers program it. "We used AI to design the chip, and designed the chip so AI could program it," the company said in its blog post.

Ho told reporters OpenAI plans to deploy Jalapeño "in small volumes" by the end of 2026, with more meaningful deployment coming in 2027. OpenAI has not disclosed how many chips it plans to ship.

Not a break with Nvidia

Despite the favorable benchmarks, Ho was clear that OpenAI isn't ditching its current suppliers. He said OpenAI's compute strategy still includes "very good partners," Nvidia among them, and that the company has no plan to replace its entire chip lineup with Jalapeño. OpenAI is already developing second- and third-generation versions of the chip.

A broader pattern flagged by financial newsletter Newsquawk shows that in-house chip programs at major AI labs tend to follow the same script: a custom chip gets announced with strong efficiency claims, gets covered as a threat to incumbent suppliers like Nvidia, and then ends up supplementing rather than replacing merchant GPU orders as deployment actually scales. Newsquawk pointed out that efficiency-focused inference chips address a different problem than frontier training hardware, and that claims made at announcement stage rarely come with independently benchmarked numbers. Jalapeño's rollout partially closes that gap given SemiAnalysis's hands-on testing but doesn't fully resolve it given the small initial deployment volumes.

Independent benchmarks back up OpenAI's speed and efficiency claims against today's best Nvidia hardware. But "small volumes" by the end of 2026 is a long way from displacing Nvidia's order book. Nvidia's own next-generation Rubin architecture is the real test still to come, and neither the SemiAnalysis benchmarks nor OpenAI's blog post address how Jalapeño stacks up against chips that aren't on the market yet. The next real checkpoint is 2027, when OpenAI says deployment volume actually ramps up.

Sources used for this briefing

This briefing was written by UBH's AI agent — these are the reporting inputs it draws on, linked so you can verify.

center-left
TechCrunchOpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks show
center-left
tech.yahooOpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks show
left
The VergeOpenAI says its Jalapeño chip can power faster AI responses than the competition
unknown
The New StackOpenAI built a chip in nine months. Then it let AI rewrite the code.
unknown
OpenAIJalapeño’s first results show industry-leading speed and efficiency in AI inference
unknown
newsletter.semianalysisOpenAI’ Jalapeño: Better Than Nvidia Blackwell
unknown
NewsquawkOpenAI's Jalapeno chip can conclude tasks more efficiently and returns responses faster than other AI systems