OpenAI has spent the past year buying Nvidia hardware by the gigawatt. This week it showed off the chip it hopes will let it buy a little less of it. On August 25 the company published the first performance figures for Jalapeno, the custom inference processor it designed with Broadcom, and the numbers are aimed squarely at the market leader.

Jalapeno is an application-specific chip built for one job: running large language models after they have been trained, the workload known as inference. On the InferenceX benchmark suite run by SemiAnalysis, a Jalapeno system delivered between 1.5 and 1.9 times more throughput per kilowatt than Nvidia's GB200 and GB300 rack systems, and cut end-to-end latency by somewhere between 1.7 and 3.6 times. The tests used three openly available models: OpenAI's own GPT-OSS 120B, DeepSeek R1, and Moonshot's trillion-parameter Kimi K2.5.

The efficiency claim rests on power. Tom's Hardware reports the chip is a single reticle-sized die built on TSMC's N3P process, rated at 700 watts but drawing 550 or less during the test runs. That is roughly half the appetite of Nvidia's flagship rack parts. Samsung is believed to supply the HBM4 memory, while Broadcom handles the silicon implementation and networking. The design went from early schematics to fabrication in about nine months, a pace OpenAI credits partly to using its own earlier models to help with the engineering.

A caveat worth keeping

These are OpenAI's own benchmarks, on a suite it chose, under conditions it set. Independent labs have not verified them, and inference results shift a great deal with batch size, model choice, and software tuning. The figures are a claim rather than a settled fact. Even so, the direction was convincing enough that Broadcom's shares rose on the news.

What makes Jalapeno matter is less the raw speed than what it signals. Nvidia's grip on the market rests partly on CUDA, the software layer that has made its GPUs the default choice for years. A chip that only has to run one company's models does not need to win over the whole industry, and that narrows the moat. OpenAI has said it will deploy Jalapeno in its own data centres later this year, with Microsoft expected to take a large slice of the first production run.

The real driver is cost. OpenAI is one of the largest buyers of AI accelerators on the planet, and every watt it saves at inference time is margin it keeps. Its own chip does not free it from Nvidia, which it still relies on for training. But it changes the shape of the negotiation, and it hands the company a hedge against the price of the hardware everyone else is scrambling to secure.

Sources

  1. i. www.tomshardware.com
  2. ii. techcrunch.com
  3. iii. newsletter.semianalysis.com
  4. iv. www.trendforce.com

Commentarii · 0

Add · a · Comment