Olix triples its valuation to $3.3B, betting the inference bottleneck is silicon, not models

Olix triples its valuation to $3.3B, betting the inference bottleneck is silicon, not models

Good morning ๐Ÿ‘‹ A chip startup just tripled its own price tag in six months, and the first customer won't see silicon for another year and a half. Below: why investors are buying the roadmap anyway.

In today's issue:

Get tomorrow's issue in your inbox.

One concise AI brief, sent after the signal clears the noise.


๐Ÿ”ญ THE ONE THING

โšก Olix triples its valuation to $3.3B, betting the inference bottleneck is silicon, not models

Olix raised $312M at a $3.3B valuation, up from just over $1B in February, with Fundomo leading and Arm, Hudson River Trading, and Reed Hastings writing checks alongside existing backers doubling down. The pitch is DX-1, a decode accelerator that claims over 10,000 tokens per second per user on 100B-parameter models using on-chip SRAM instead of the HBM everyone else is scrapping over. That's a real architectural bet, not a spec-sheet flex, which is probably why Arm and a high-frequency trading shop both wanted in. First customers don't see silicon until H2 2027, so this valuation is built entirely on a roadmap: six months, 3x the price tag, zero shipped chips. I want benchmarks against actual Nvidia hardware before I call this the inference chip.


๐Ÿง  MODELS & RELEASES


๐Ÿ”ฌ RESEARCH HIGHLIGHTS


๐Ÿ› ๏ธ TRY THIS

Put a number on the local-vs-API question, using this week's tool

1. Clone debpalash/OmniVoice-Studio and point it at whatever GPU you already have, even a rented spot instance if you don't own one.

2. Run the same job you'd normally send to ElevenLabs: clone a voice, dub a short clip, or dictate a paragraph.

3. Time it. Log $/minute against your last ElevenLabs invoice for that same task.

4. If local wins on cost and the quality holds, move the recurring work off the API. If it doesn't, at least you're arguing from a number instead of a hunch, which matters while everyone from Olix on down is repricing what inference actually costs.

Prompt: Write a Python script that times OmniVoice-Studio's voice-clone inference on a sample WAV file, tracks GPU memory used, and outputs a $/minute figure given my hourly compute cost.

Worth a look


๐Ÿ”— QUICK LINKS


That's the lineup. Back tomorrow.

Pradeep Perugu

Get inovAIte in your inbox.