SoftBank closes its $64.6B OpenAI bet, almost all of it borrowed
Good morning ๐ SoftBank just finished writing a $64.6 billion check to OpenAI. Almost none of it was SoftBank's own money.
In today's issue:
- ๐ญ SoftBank closes its $64.6B OpenAI bet, almost all of it borrowed
- ๐ง Cloudflare's instant decision models, GPT-6 Astra lands on Blackwell
- ๐ฌ Agents learning what to forget, and a benchmark that credited tool calls that never happened
- ๐ Halluminate's $30M Series A for finance-agent training data
- ๐๏ธ A $300M GPU-smuggling indictment
- ๐ ๏ธ An hour-long audit for your agent's citations
- ๐ Amazon's chip sale-leaseback, AWS's town money, and ElevenLabs' $22B tender
Get tomorrow's issue in your inbox.
One concise AI brief, sent after the signal clears the noise.
๐ญ THE ONE THING
๐ฐ SoftBank closes its $64.6B OpenAI bet, almost all of it borrowed
SoftBank wired the final $10 billion tranche of its OpenAI commitment on October 1, Japan time (its own release has the exact figure: ยฅ1,579.6 billion), closing out the three-part, $30 billion follow-on it announced back in February. That puts SoftBank's cumulative OpenAI stake at $64.6 billion for roughly 13% of the company, a number GuruFocus and half a dozen other outlets all confirm without dispute. None of that $10 billion came from SoftBank's cash pile: it's funded by $11.1 billion in senior notes priced September 24, carrying coupons up to 9.75%, which also let SoftBank cancel the untouched $10 billion still sitting on its OpenAI bridge facility. Do the arithmetic and $64.6 billion for 13% implies an average entry price south of $500 billion, well under the $852 billion post-money mark OpenAI's own March round set, so SoftBank is sitting on a large paper gain at today's price and a debt-service bill either way. Watch the second number, not the first: a conglomerate borrowing at nearly 10% to hold a bigger slice of an unprofitable lab is the real story here.
๐ง MODELS & RELEASES
- ๐ Cloudflare shipped Clef and Clef-flash, open-weight decision models (Apache 2.0, 27B and 9B) that skip text generation and just return probabilities over a set of typed answers. Flash hits 38.8ms median latency, 13x faster than Jev's 524.1ms, and the pair wins 7 of 10 decision benchmarks outright. If you're doing classification or routing at the edge, this is the cheaper, faster option you've been overpaying to avoid.
- ๐งฎ Allen Institute for AI open-sourced Olmo-core 3, the training infrastructure behind its push past a trillion parameters in mixture-of-experts models. Swapping FSDP for DDP and keeping experts resident on GPU got them 2.7x the throughput, 52,000 tokens per second per GPU on B300s. The code's public, which matters more than the benchmark: a lab without hyperscaler budget can now train something MoE-shaped that doesn't collapse at scale.
- โก NVIDIA and OpenAI put GPT-6 Astra Ultrafast on Blackwell silicon, claiming up to 8x faster inference in the API and for ChatGPT Work and Codex users. Treat the 8x as a vendor number until someone outside the two companies reproduces it. Even half of that moves the cost math on anything latency-sensitive.
๐ฌ RESEARCH HIGHLIGHTS
- AutoCompact trains coding agents to decide, as part of their own policy, when to compress a long trajectory and what to keep. The payoff: 9.2 points absolute on SWE-bench Verified over the base agent, 5.0 on SWE-PolyBench Verified, holding even in a 16K window that has to fall back on compaction mid-task. Builder read: the ceiling on long agent runs isn't context length, it's agents that don't know what to throw away. link
- SFT might not be the weak sibling after all. Standard story: RL generalizes, supervised finetuning forgets and overfits. A new paper resamples off-policy expert data (an MCMC pass that nudges it closer to what the model would generate itself) before running plain SFT, and the result rivals RL baselines on math and science tasks, forgetting less in the process. Watch, don't act: one paper, strong claims, worth a replication before it changes your posttraining stack. link
- A 661.6M-parameter model just embarrassed its bigger sibling. Two similarly-built Spanish tool-use models scored nearly identically on a lenient keyword benchmark (0.660 vs. 0.650). A cheap verbatim-reproduction check split them apart instantly: the small model nailed 6 of 6 held-out tool calls, the 1.1B model nailed zero, because its web-heavy training had wiped out the probability of ever emitting the tool-call token. Builder read: if your eval is keyword matching, you may be crediting tool use that never happened. link
๐ AI STARTUPS
- Halluminate raised a $30M Series A, led by Oak HC/FT with Y Combinator and FT Partners back again, to build the RL environments and benchmarks that teach agents real finance work instead of toy tasks. Ten months, zero to a mid-eight-figure revenue run rate, profitable the whole way, and already building environments for four of the five major closed-source labs. That's the quiet infrastructure bet: whoever owns the training data for "diligence done right" ends up deciding how good finance agents actually get. Series A announcement
๐๏ธ POLICY & REGULATION
- DOJ indicted Greg Lui, a San Gabriel server reseller, for routing more than $300 million in US-made GPU servers to China, laundering the shipments through Malaysia and Singapore to dodge export licenses and faking end-user paperwork along the way. His company, Earthmade Computer, pulled in $176 million from the scheme over two years before anyone caught it. Conspiracy and money-laundering counts each carry up to 20 years. DOJ press release
๐ ๏ธ TRY THIS
$64.6B is now riding on agents actually doing the job. Spend an hour checking yours cites the right source, not just a true fact.
1. Pull the last 20 answers your RAG or MCP agent gave with a citation attached.
2. Break each one into individual claims, and note which source the agent pointed to for each.
3. Re-check every claim against *only* its cited document, not your full context window. A fact can be true somewhere in your corpus and still be attributed to the wrong file.
4. Tally the misses. The technique below caught 138 of 139 unsupported claims this way, and flagged all 50 deliberate source swaps a test set threw at it. A rough version by hand will still catch your obvious offenders.
Prompt: Check this claim only against the source text given, ignore anything else you know. Claim: "{{claim}}". Cited source: {{source_name}}. Source text: {{source_text}}. Reply: supported, unsupported, or wrong-source (true, but not found in this document).Worth a look
- Multiverse Computing's ProvenanceGuard is the source for that technique: a post-generation check that keeps each claim tied to a specific source instead of pooling everything into one blob, which is how agents end up citing a real fact from the wrong document. Source-Aware Verification for MCP Agents
- AWS and NVIDIA show the other half of the problem: agents that actually remember. Amazon S3 Vectors plugs into the NeMo Agent Toolkit as a memory store that scales to 2 billion vectors without you provisioning anything. Build agent memory with NeMo and S3 Vectors
๐ QUICK LINKS
- Amazon is shopping $8B of its Nvidia Grace Blackwell chips to outside investors through a special-purpose vehicle, then leasing the same racks straight back so the depreciating silicon comes off its balance sheet. FT
- AWS is putting $1B over five years into the towns hosting its data centers (free community college, efficiency grants, local funding pots) as moratorium talk spreads across more than 100 places. Amazon
- ElevenLabs closed a $300M employee tender that doubles its valuation to $22B, with Goldman Sachs and GIC joining as new backers alongside existing investors like a16z. ElevenLabs
That bet only works if OpenAI's growth outruns SoftBank's interest bill.
Pradeep Perugu