OpenAI says its new Jalapeno chip can make AI faster while using less power

OpenAI unveiled its new Jalapeno AI chip at a recent conference. This processor significantly boosts AI response speeds and reduces power consumption. The chip delivers improved performance per watt and lower end-to-end latency. OpenAI plans to...

ANI
"We made a chip and it is fast": OpenAI unveils Jalapeno custom AI inference chip
OpenAI has unveiled new performance results for its custom AI chip, Jalapeno, at the Hot Chips conference in Palo Alto, California, saying the processor can make AI responses faster while using less power.

The company said Jalapeno delivered between 1.5 and 1.9 times more AI work for every unit of power used, while reducing the time taken to generate a response by between 1.7 and 3.6 times across three AI models it tested.

For highly interactive workloads, performance was between 2.1 and 4.1 times higher, OpenAI said.


The announcement is part of OpenAI's push to build more of the computing infrastructure needed to run its AI services in-house, as demand for AI continues to drive up the need for expensive chips and electricity.

OpenAI CEO Sam Altman summed up the announcement rather more simply on X, posting: "we made a chip and it is fast."

<blockquote class="twitter-tweet"><p lang="en" dir="ltr">we made a chip and it is fast</p>— Sam Altman (@sama) <a href="https://x.com/sama/status/2092339694210040187?ref_src=twsrc%5Etfw">August 25, 2026</a></blockquote> <script async="" src="https://platform.x.com/widgets.js" charset="utf-8"></script>
What is Jalapeno?
ADVERTISEMENT

Jalapeno is an inference chip. In simple terms, it is designed to run AI models and generate answers after those models have already been trained.

Think of AI training as teaching a student, while inference is the student actually answering the questions. Jalapeno is built for the second part.

That means the chip is intended for the enormous amount of computing that happens when users ask ChatGPT questions or when AI agents carry out tasks.

OpenAI said it designed Jalapeno specifically around the way modern AI models work, rather than adapting a general-purpose chip for the job.
ADVERTISEMENT

Also Read: India’s AI data-centre push could drive 5% of global chip demand by 2030

Faster AI with less power
ADVERTISEMENT

One of the biggest challenges in running AI services is balancing speed with efficiency.

A system can be designed to handle a large number of requests at once, but that does not necessarily mean each user gets a faster response. Optimising for very low response times can also require more computing power.

OpenAI says Jalapeno is designed to improve both.

The company tested the chip on three publicly available models — GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T — using InferenceX, a benchmark from SemiAnalysis that measures how quickly and efficiently AI systems serve requests.

On Kimi K2.5 1T, the largest model tested, OpenAI said Jalapeno delivered about 1.5 times higher performance per watt and 3.4 times lower end-to-end latency than the comparison system.

In plain English, that means the chip could do more AI work using the same amount of electricity, while also getting answers to users faster.

Jalapeno has a rated power consumption of 700 watts, although OpenAI said its measured sustained power stayed at or below 550 watts during the workloads it tested.

Why does this matter?

Every ChatGPT request requires computing power. As more people use AI — and as AI agents perform longer, more complicated tasks — the amount of computing required continues to increase.

If OpenAI can get more AI work from the same amount of electricity and hardware, it could serve more users without infrastructure costs rising at the same rate.

The chip has also been designed to reduce the amount of data that has to move between different parts of a computer system. This matters because AI models constantly move large amounts of information between processors and memory while generating responses.

Also Read: Google expands Gemini AI platform for law firms, lawyers

OpenAI said Jalapeno is designed to keep important model information closer to the computing resources that need it, reducing delays.

AI helped build the AI chip

There is another unusual aspect to Jalapeno. OpenAI says its own AI models helped develop the chip.

The company said AI helped it move from the initial design to "tapeout" — the stage when a chip design is finalised for manufacturing — in nine months.

AI was used to explore different chip designs, speed up testing and verification, and optimise parts of the processor.

OpenAI also used its AI coding tools to program Jalapeno for models that were not originally part of its production plans. The company said AI-generated implementations for some parts of GPT-OSS ran 1.5 to 1.8 times faster than versions written by human experts.

OpenAI will still use Nvidia

Despite developing its own chip, OpenAI said it will continue to use accelerators from Nvidia and other partners for both training and inference.

Jalapeño is therefore not an immediate replacement for Nvidia hardware. Instead, it gives OpenAI another source of computing power and greater control over the infrastructure needed to run its AI models.

OpenAI plans to begin deploying Jalapeno in its own computing infrastructure by the end of 2026. The company said a second-generation chip is already deep in development, while a third generation is also taking shape.
Download
The Economic Times Business News App
for the Latest News in Business, Sensex, Stock Market Updates & More.
READ MORE
ADVERTISEMENT

READ MORE:

LOGIN & CLAIM

50 TIMESPOINTS

More from our Partners

Loading next story
Business News › AI › AI Insights › OpenAI says its new Jalapeno chip can make AI faster while using less power
Text Size:AAA
Success
This article has been saved

*

+