OpenAI's first chip beats Nvidia at its own game, on paper

OpenAI's first chip beats Nvidia at its own game, on paper

Good morning πŸ‘‹ OpenAI just put its own inference chip on the record against Nvidia's best, and the numbers are the kind that make you want to see someone else run them.

In today's issue:

Get tomorrow's issue in your inbox.

One concise AI brief, sent after the signal clears the noise.


πŸ”­ THE ONE THING

🌢️ OpenAI's first chip beats Nvidia at its own game, on paper

OpenAI put real numbers behind JalapeΓ±o this week: its first inference ASIC, built with Broadcom, pulled 1.5 to 1.9x the throughput per kilowatt and up to 3.6x lower latency than Nvidia's GB300 racks on the InferenceX benchmark, presented alongside a Hot Chips 2026 talk. That's a 700W part outrunning Nvidia's 1,400W flagship. Read the fine print before you touch your Nvidia allocation: the test excludes speculative decoding and pits JalapeΓ±o running single-token against GB300 running multi-token, and no one outside OpenAI has put the chip on a bench yet. Still, "OpenAI built its own inference silicon and it's not embarrassing" is the headline that actually matters here. Small volumes land by end of 2026; the real verdict comes in 2027, once it has to perform at scale on someone else's stopwatch.


🧠 MODELS & RELEASES


πŸ”¬ RESEARCH HIGHLIGHTS


πŸ› οΈ TRY THIS

Stress-test a vendor's "early results" claim before you plan around it

1. Pull the actual benchmark writeup OpenAI published for the chip, not the recap, and note the workload specifics: batch size, sequence length, model, precision.

2. Re-run a workload of the same shape on your current inference stack. Log tokens/sec and cost per million tokens.

3. Put the two numbers side by side. A gap past 2x is a sign you're comparing OpenAI's best case to your average case, not the chip.

4. Hold off on infra commitments until someone with no stake in the outcome reproduces the number. Day-one vendor benchmarks are marketing until then.

Prompt: Given this benchmark writeup [paste text], list every workload parameter (batch size, sequence length, precision, model) it doesn't disclose, and flag which omissions would most change the reported numbers.

πŸ”— QUICK LINKS


See you tomorrow.

Pradeep Perugu

Get inovAIte in your inbox.