OpenAI Jalapeño: The Inference Chip That Actually Beats Blackwell

There's a narrative everyone accepted: first-gen custom chips never compete with Nvidia. OpenAI just torched that narrative with a blowtorch.

Today, SemiAnalysis published a deep-dive after OpenAI invited them into the lab to benchmark Jalapeño — OpenAI's first inference chip, built from scratch in partnership with Broadcom. The headline is simple: Jalapeño beats Nvidia Blackwell on perf/W across almost every scenario tested.

The numbers are not soft. SemiAnalysis ran their InferenceX suite in-person with OpenAI engineers. On DeepSeek R1 at concurrency 1, Jalapeño hits 700+ tokens/sec per user. On the broader InferenceX benchmark suite, it consistently outruns Blackwell — without Multi Token Prediction (MTP), while the Blackwell numbers include MTP.

Key architectural details:

The honest caveat: Jalapeño is still an engineering sample. SemiAnalysis notes the fair comparison is really against Nvidia Vera Rubin (which ships to customers right now, using HBM4 too), not Blackwell. Rubin's NVL72 delivers 5.4× the perf/MW of GB200 NVL72. But even positioning Jalapeño against Rubin, a custom chip beating Nvidia's best in its own architecture-optimized sweet spot is a statement.

This matters for a few reasons:

Bottom line: the "first-gen custom chip" excuse just evaporated. Jalapeño proves you can go from blank slate to industry-leading inference in one shot — if you have the talent, the capital, and the vertical integration. Nvidia's response will be instructive.