OpenAI's first chip beats Nvidia at its own game, on paper
Good morning π OpenAI just put its own inference chip on the record against Nvidia's best, and the numbers are the kind that make you want to see someone else run them.
In today's issue:
- π OpenAI's first chip beats Nvidia at its own game, on paper
- π§ A 4-bit model that beats its own full-precision parent
- π¬ Decorative chain-of-thought, a confidently wrong chess-playing model, and where self-improving agents stall
- π οΈ Try this: stress-test a vendor's benchmark before you plan around it
- π OpenAI's Russia takedown and Admin plugin, MiniMax's revenue jump, SoftBank's bond plan
Get tomorrow's issue in your inbox.
One concise AI brief, sent after the signal clears the noise.
π THE ONE THING
πΆοΈ OpenAI's first chip beats Nvidia at its own game, on paper
OpenAI put real numbers behind JalapeΓ±o this week: its first inference ASIC, built with Broadcom, pulled 1.5 to 1.9x the throughput per kilowatt and up to 3.6x lower latency than Nvidia's GB300 racks on the InferenceX benchmark, presented alongside a Hot Chips 2026 talk. That's a 700W part outrunning Nvidia's 1,400W flagship. Read the fine print before you touch your Nvidia allocation: the test excludes speculative decoding and pits JalapeΓ±o running single-token against GB300 running multi-token, and no one outside OpenAI has put the chip on a bench yet. Still, "OpenAI built its own inference silicon and it's not embarrassing" is the headline that actually matters here. Small volumes land by end of 2026; the real verdict comes in 2027, once it has to perform at scale on someone else's stopwatch.
π§ MODELS & RELEASES
- ποΈ Multiverse Computing compressed GPT-OSS 120B down to a 4-bit, 60B-parameter model using a new distillation method called Quantization-Aware Healing, and the shrunk version beat its full-precision 60B counterpart on 7 of 9 benchmarks (long-context reasoning up 7.4 points, AIME math up 5.6). It also converged in roughly 100 training steps versus 700 for standard quantization-aware training. Quantization-Aware Healing
π¬ RESEARCH HIGHLIGHTS
- The reasoning is decorative. A perturbation audit ran 14 medical LLMs through 30 kinds of edits to their chain-of-thought (flipped severity, swapped demographics, ablated evidence) and checked whether the visible reasoning actually tracked the final answer. 73% of the time it didn't: corrupt the chain, the diagnosis stays put. Two board-certified clinicians confirmed the edits were clinically valid, so this isn't a measurement artifact. If you're using CoT as an audit trail for a clinical model, you're auditing theater. arXiv
- Confident and wrong. Researchers built a hidden-information chess variant, where a king's identity can secretly relocate, and asked models to state confidence separately from their move. When a model was at least 50% sure it had found the hidden king, it was right once in 62 tries. The uglier finding: the setup that looked best on legality, latency, and cost had the worst-calibrated beliefs of the group. Gating agent actions on the model's own stated confidence is gating on noise. arXiv
- Meta^n goes after why self-improving agents stall. Add a level that reflects on the agent's process, and something in that loop has to stay frozen to keep it stable, usually capping real gains at meta-depth two. This paper fixes the meta-operation itself instead of the content it edits, then lets depth emerge from convergence rather than a preset limit. It's the only method in the paper's eight-benchmark suite to score above zero on ARC-AGI-2, a benchmark built specifically to resist memorized tricks. One paper, one lab, worth tracking before calling it solved. arXiv
π οΈ TRY THIS
Stress-test a vendor's "early results" claim before you plan around it
1. Pull the actual benchmark writeup OpenAI published for the chip, not the recap, and note the workload specifics: batch size, sequence length, model, precision.
2. Re-run a workload of the same shape on your current inference stack. Log tokens/sec and cost per million tokens.
3. Put the two numbers side by side. A gap past 2x is a sign you're comparing OpenAI's best case to your average case, not the chip.
4. Hold off on infra commitments until someone with no stake in the outcome reproduces the number. Day-one vendor benchmarks are marketing until then.
Prompt: Given this benchmark writeup [paste text], list every workload parameter (batch size, sequence length, precision, model) it doesn't disclose, and flag which omissions would most change the reported numbers.π QUICK LINKS
- OpenAI banned a Russia-linked network that used ChatGPT to run a fake Israeli think tank pushing a "sovereignty index" flattering Moscow and knocking the West. details
- OpenAI shipped an Admin plugin for ChatGPT Work and Codex. Usage analytics, seat permissions, limits, one panel. plugin
- Hugging Face published a guide to wiring multi-step AI workflows in Gradio, handy if you're stitching together agent pipelines without a heavier framework. guide
- MiniMax posted H1 revenue up 283% year over year to roughly $116.6M, outpacing 2025's 159% clip. Its M3 model still trails Z.ai's GLM-5.2. report
- SoftBank is weighing a $10-20B bond sale, dollars and euros, to chip away at the $40B bridge loan behind its OpenAI stake. Pricing could land as early as September. filing
See you tomorrow.
Pradeep Perugu