OpenAI Jalapeño: The Inference Chip That Actually Beats Blackwell
There's a narrative everyone accepted: first-gen custom chips never compete with Nvidia. OpenAI just torched that narrative with a blowtorch.
Today, SemiAnalysis published a deep-dive after OpenAI invited them into the lab to benchmark Jalapeño — OpenAI's first inference chip, built from scratch in partnership with Broadcom. The headline is simple: Jalapeño beats Nvidia Blackwell on perf/W across almost every scenario tested.
The numbers are not soft. SemiAnalysis ran their InferenceX suite in-person with OpenAI engineers. On DeepSeek R1 at concurrency 1, Jalapeño hits 700+ tokens/sec per user. On the broader InferenceX benchmark suite, it consistently outruns Blackwell — without Multi Token Prediction (MTP), while the Blackwell numbers include MTP.
Key architectural details:
- HBM4 memory — comparable to flagship GPUs from Nvidia and AMD
- Generalized inference chip, not over-specialized for OpenAI's own models
- Extreme hardware-software co-design — the chip and compiler were built together
- As a joke, they ported Doom to Jalapeño using just Codex prompts
The honest caveat: Jalapeño is still an engineering sample. SemiAnalysis notes the fair comparison is really against Nvidia Vera Rubin (which ships to customers right now, using HBM4 too), not Blackwell. Rubin's NVL72 delivers 5.4× the perf/MW of GB200 NVL72. But even positioning Jalapeño against Rubin, a custom chip beating Nvidia's best in its own architecture-optimized sweet spot is a statement.
This matters for a few reasons:
- The CUDA moat just got a stress test. If OpenAI can build a competitive inference chip in one generation, others can too. The argument that "nobody can beat Nvidia at their own game" is weaker today.
- AI chip design with AI — OpenAI claims Jalapeño's design was accelerated by AI tools, and the Doom-as-Codex-demo is a flex that the software stack is mature enough for agent-driven bringup.
- Inference is the battlefield. Training still happens on clusters, but inference is where the economics break for everyone deploying models at scale. A chip that beats Blackwell on perf/W shifts the TCO calculus for any serious deployment.
Bottom line: the "first-gen custom chip" excuse just evaporated. Jalapeño proves you can go from blank slate to industry-leading inference in one shot — if you have the talent, the capital, and the vertical integration. Nvidia's response will be instructive.