OpenAI’s Jalapeño chip is built for fast inference at scale

OpenAI and Broadcom unveiled Jalapeño, a custom AI inference chip that beats Nvidia Blackwell on speed and efficiency at 700W vs 1,200-1,400W on the public InferenceX benchmark.

OpenAI’s Jalapeño chip is built for fast inference at scale

OpenAI just pulled back the curtain on its first custom AI inference chip, and the benchmark lead is worth your attention. Jalapeño, built with Broadcom, delivers faster responses and better efficiency than the Nvidia Blackwell systems it was tested against.

A custom chip is a signal that the biggest AI labs aren’t content to buy their most important hardware off the shelf anymore.

I haven’t touched the silicon myself, but the benchmark data OpenAI published on August 25, 2026 is specific enough to matter for anyone who uses ChatGPT, Codex, or the API OpenAI’s Jalapeño first results. This is research-based coverage of a hardware launch, not a hands-on review.

The chip race is heating up across the industry, and OpenAI is the last of the frontier labs to confirm its own silicon.

What is an AI inference chip anyway

An AI inference chip does the work of actually running a trained model after it’s built. When you type a prompt and ChatGPT answers, that answer is produced by inference hardware, not by the training clusters that cost billions to build OpenAI and Broadcom unveil Jalapeño.

Training gets the headlines, but inference is where the bills actually arrive every time someone sends a prompt.

Inference is where AI actually reaches people, so every gain in speed or cost shows up as a snappier ChatGPT reply or a cheaper API call. That’s why OpenAI is now designing its own accelerator instead of relying only on Nvidia and other partners. Owning the accelerator means OpenAI can tune the whole path from model to answer.

The benchmark that has everyone talking

Content-only crop of the official OpenAI Jalapeño first results blog post body at openai.com/index/jalapeno-first-results/. The visible top text is the section heading 'Jalapeño widens the lead at previous-best TBT' followed by the chart sub-heading 'Mixed-token throughput matched at the existing accelerator's fastest decoding speed'. Below that sits a horizontal legend with a green filled circle labeled 'Jalapeño' and a blue filled circle labeled 'Existing best'. The chart itself is a vertical bar chart with two grouped pairs of bars on a dark background. The Y-axis label reads 'Mixed tokens / second / kW' with tick marks 0, 5000, 10000, 15000, 20000, 25000. The X-axis shows two category groups: 'GPT-OSS - 535.28 tokens/s/user' and 'DeepSeek R1 - tokens/s/user' (partially cut off at the right edge). The leftmost Jalapeño green bar is labeled '22,935' and the Existing-best blue bar is labeled '427', with a pill-shaped callout reading '53.7x' and a vertical arrow spanning the pair. The second Jalapeño green bar is labeled '12,258'. No global OpenAI nav, footer, breadcrumbs, sign-in overlay, or other chrome is included.
OpenAI's published InferenceX benchmark chart: Jalapeño vs the existing-best accelerator on mixed-token throughput per kilowatt (GPT-OSS 120B and DeepSeek R1). Image: openai.com/index/jalapeno-first-results/ (first-party body asset).

OpenAI tested Jalapeño on InferenceX, a public benchmark from SemiAnalysis, against Nvidia’s GB200 and GB300 superchips. Across GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 1T, this AI inference chip delivered 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency than the comparison systems.

For highly interactive workloads like agent tasks, the chip posted 2.1 to 4.1 times higher performance than the commercial systems in the comparison. That matters because agents run many steps in sequence, so small delays compound into a slow overall task.

Jalapeño is rated at 700 watts, though OpenAI says its measured sustained power stayed at or below 550 watts on the tested workloads. The comparison Nvidia systems are rated at 1,200 to 1,400 watts, which is a big part of why the per-watt numbers favor this AI inference chip.

Why this is a big deal for OpenAI users

If the efficiency holds up at scale, the practical upside is simple: faster responses, more responsive agents, and more reliable access when demand spikes. OpenAI says the gains help make increasingly capable AI more affordable and more broadly available.

Cheaper inference is the quiet engine behind every free tier and price cut you’ve used this year.

This also strengthens the case for coding agents, which is why our best AI coding agents roundup keeps tracking inference speed as a differentiator. A chip that cuts latency on multi-step agent runs changes what those tools feel like in daily use.

The full-stack strategy

Jalapeño isn’t a one-off. OpenAI describes it as the first step in a multigenerational compute platform, with Gen 2 deep in development and Gen 3 taking shape. The company says it moved from initial design to manufacturing tape-out in just nine months, accelerated by its own AI models.

OpenAI also says AI-generated code ran 1.5 to 1.8 times faster than human-expert implementations on selected blocks when the team brought open-weight models to the chip. That’s a hint about where chip design itself is heading: the models help build the hardware that runs the models.

If the models can design the chips, the loop gets faster with every generation.

When you will actually see it

Don’t expect Jalapeño in your ChatGPT tomorrow. OpenAI plans to begin deploying it within its own compute infrastructure by the end of 2026, with more significant deployment in 2027. Richard Ho, OpenAI’s head of hardware, said it would deploy at the end of 2026 in very small volumes.

The comparison against Nvidia is also a moving target, since Blackwell won’t stand still while Jalapeño ramps. But the direction is clear: OpenAI wants to own more of the stack underneath the models you already use.

Price cuts on models like GPT-5.6 show the other half of this story, because cheaper inference is what makes those cuts possible OpenAI price cuts for GPT-5.6 Luna and Terra. Jalapeño is the hardware bet behind that trend.

The bottom line

OpenAI’s Jalapeño is the most concrete sign yet that the company is building the full stack, from models down to the chip that serves them. The benchmark lead over Nvidia Blackwell is real on paper, and if it translates to production, your next ChatGPT answer could arrive a little faster and cost OpenAI a little less.

The labs that control their own silicon will set the pace for everyone else.

Tony Simons

Reviewed & Written By

Tony Simons

Independent tech reviewer and creator of Tony Reviews Things. 14 years of hands-on testing, software auditing, and workflow automation. I test the gear so you don't waste your money on junk.

Submit a Take

Your email address will not be published. Required fields are marked *