OpenAI’s Jalapeno chip takes aim at Nvidia

OpenAI is making a bigger move into the chip business, and this time it is not just about having its own hardware. It is about cutting one of the biggest costs in AI: running models at scale.

The company says its new custom chip, called Jalapeno, has outperformed Nvidia’s GB300 in certain tests focused on AI inference. That matters because inference is the part of AI where the real-world work happens after a model has been trained. Every time someone asks ChatGPT a question, generates an image or runs an AI task, computing power is being used.

If OpenAI can make those responses faster and cheaper with its own chips, the economics of running AI services could change significantly.

OpenAI is taking a shot at Nvidia’s strongest position

Nvidia remains the dominant supplier of AI processors, but OpenAI clearly does not want to depend entirely on one chipmaker as its computing needs explode.

Jalapeno was tested against Nvidia’s GB300, which was the leading processor on the public benchmarking system used for the comparison.

According to OpenAI, Jalapeno came out ahead in two important areas:

  • Performance per unit of power
  • Speed of delivering AI responses

Those are two metrics that matter enormously for an AI company operating at massive scale.

A chip that can deliver more AI work while consuming less electricity can reduce data center costs. A chip that responds faster can improve the experience for customers who care about latency.

That gives OpenAI something potentially more valuable than simply claiming that its chip is “faster.”

It gives the company another way to control the economics of AI.

The real opportunity is inference

There is an important distinction here.

Jalapeno is not designed to replace Nvidia across the entire AI stack.

The chip is aimed at inference, rather than training AI models.

Training is the stage where models learn from enormous amounts of data. Inference is what happens afterward, when those trained models are actually responding to users and performing tasks.

That distinction is important because Nvidia remains extremely strong in AI training.

OpenAI is instead going after the part of the business that could become an enormous ongoing expense as AI adoption grows.

Think about it this way: training a model is a major upfront computing challenge. But once millions of people start using that model, inference becomes a constant cost.

Every prompt has a price.

Every response has a price.

Every additional user adds more computing demand.

That is where a more efficient chip could make a meaningful difference.

Why power consumption matters so much

AI chips are incredibly power hungry, and data centers are becoming some of the biggest consumers of electricity.

OpenAI says Jalapeno can deliver strong performance at 700 watts, which could help reduce the cost of running its data centers.

That may sound like a technical detail, but it is actually central to the business case.

AI infrastructure is not just about buying chips.

Companies have to pay for:

  • Electricity
  • Data center capacity
  • Cooling
  • Networking
  • Storage
  • Hardware maintenance
  • The chips themselves

When you operate AI services at enormous scale, even a relatively small improvement in efficiency can translate into substantial savings.

That is why OpenAI developing its own processor is about much more than competing with Nvidia on a benchmark chart.

It is ultimately a cost-control strategy.

Jalapeno is trying to solve two problems at once

One of the more interesting claims from OpenAI is that Jalapeno can perform well in both high-throughput and low-latency workloads.

High throughput means the chip can handle a large amount of AI work efficiently.

Low latency means users get answers quickly.

Usually, companies have to make tradeoffs between these two goals. OpenAI says Jalapeno is designed to perform well in both areas.

That could give OpenAI more flexibility in deciding which models should run on its own hardware.

For some customers, the priority could be lower cost.

For others, it could be faster responses.

The same underlying chip could potentially support both use cases.

But this is not Nvidia’s latest generation

There is an important caveat to the headline.

OpenAI tested Jalapeno against Nvidia’s GB300, not Nvidia’s newest Vera Rubin generation, which has just started shipping.

So the claim is not that OpenAI has built a chip that universally beats everything Nvidia makes.

It is a more specific claim: Jalapeno performed better than the GB300 in particular inference tests.

That distinction matters.

Nvidia is moving quickly, and OpenAI will have to keep improving its hardware if it wants its chips to remain competitive as Nvidia releases newer generations.

Still, OpenAI does not need to replace Nvidia entirely for Jalapeno to be valuable.

Even taking a portion of its enormous computing workload in-house could give OpenAI more control over costs and supply.

Broadcom is part of the story

OpenAI developed Jalapeno in partnership with Broadcom, which specializes in custom chips for major technology companies.

The partnership is significant because designing a competitive AI processor from scratch is an enormous undertaking.

OpenAI brings the AI workloads and software expertise.

Broadcom brings deep semiconductor and chip-design capabilities.

Together, they were able to develop Jalapeno relatively quickly, with the companies previously highlighting the speed of the development process.

And OpenAI is already looking ahead.

A second-generation chip is reportedly well into development, with the company expecting to reach the tape-out stage in the coming months.

It is also already working on concepts for a third generation.

That tells you this is not being treated as a one-off experiment.

OpenAI wants chips to become a long-term part of its infrastructure strategy.

The chip race is getting crowded

OpenAI is far from the only company trying to reduce its dependence on Nvidia.

A growing number of startups and technology companies are developing specialized AI processors.

Companies such as Cerebras, Etched and MatX are pursuing different approaches to the same fundamental problem: how do you run increasingly powerful AI models without allowing computing costs to grow out of control?

OpenAI already uses Cerebras technology for some of its models.

But OpenAI says Cerebras chips are better suited to smaller models, while Jalapeno is designed to handle larger workloads.

That does not mean OpenAI is abandoning outside suppliers.

Quite the opposite.

OpenAI’s Richard Ho made clear that the company still expects to need large quantities of Nvidia chips and other computing providers.

That makes sense.

OpenAI’s appetite for compute is enormous, and no single chip is likely to handle every workload.

OpenAI is building a multi-chip strategy

This may be the bigger story.

OpenAI is not necessarily trying to replace Nvidia.

It is trying to make sure Nvidia is not the only answer.

The company already works with multiple computing providers and chip technologies. Adding its own processor gives OpenAI another lever.

If Jalapeno is cheaper for certain inference workloads, OpenAI can shift those workloads onto its own hardware.

If Nvidia remains better suited for training or particular advanced workloads, OpenAI can continue using Nvidia.

If another specialized chip performs better for a specific model, OpenAI can use that too.

That kind of flexibility could become increasingly important as AI infrastructure becomes one of the biggest costs in the industry.

The bigger bet is on scale

OpenAI’s AI ambitions require an extraordinary amount of computing power.

As its models become larger and more capable, the company needs more chips, more data centers and more electricity.

That creates a difficult equation.

More users mean more revenue, but they also mean more computing costs.

The company therefore has to keep improving the economics of every interaction.

Jalapeno is one attempt to do exactly that.

OpenAI says the chip performed particularly well on larger and more demanding workloads in internal testing. Its public tests included one of the company’s smaller open-source models as well as models from DeepSeek and Moonshot AI.

The strongest gains were seen on Moonshot’s Kimi model, which was the largest model included in the public testing.

That could be an important signal because AI workloads are only getting bigger.

This is step one, not the finish line

OpenAI itself appears to view Jalapeno as the beginning of a much larger effort.

The company is already developing its second chip generation and thinking about a third.

That suggests the goal is not simply to build one good processor.

The goal is to build an entire hardware roadmap that evolves alongside OpenAI’s models.

And that could eventually give OpenAI something it has never fully controlled before: a larger portion of the technology stack behind its AI products.

From models to software to infrastructure to custom silicon, the company is steadily moving deeper into the machinery that makes AI possible.

What this means for Nvidia

The headline may sound like a direct challenge to Nvidia, but the relationship is more complicated.

OpenAI still needs Nvidia.

Nvidia remains a critical supplier, particularly for training and other workloads where its processors are highly competitive.

But every time OpenAI develops another workload that can run efficiently on its own silicon, its dependence on Nvidia decreases at the margin.

And if other major AI companies follow the same path, Nvidia could eventually face a more fragmented market.

That does not necessarily mean Nvidia loses its dominance.

It means the definition of dominance could change.

Instead of supplying virtually every important AI workload, Nvidia may increasingly have to compete with custom chips designed specifically for individual companies and applications.

The takeaway

Jalapeno is less about beating Nvidia and more about changing OpenAI’s economics.

The chip has not been tested against Nvidia’s newest Vera Rubin generation, and it is not designed for AI training. But within its intended role, OpenAI says it has delivered impressive results in speed and power efficiency.

If those results translate from the lab into large-scale production, the impact could be significant.

OpenAI could lower the cost of serving AI models, reduce its reliance on outside chip suppliers and gain more control over the infrastructure behind its products.

And with a second generation already in development, this looks like the beginning of a much bigger hardware push.

The AI race is no longer just about who builds the smartest model. It is increasingly about who can run those models most efficiently, at the lowest possible cost, and at massive scale.

That is the game OpenAI is now trying to play.