OpenAI's Ultrafast mode pushes GPT-5.6 Sol to 14x speed, courtesy of Cerebras silicon
Good morning ๐ The bottleneck in your agent demo was never the model. Today's lead is about the chip underneath it, and why 750 tokens a second changes what "real-time" actually means.
In today's issue:
- ๐ญ OpenAI's Ultrafast mode pushes GPT-5.6 Sol to 14x speed, courtesy of Cerebras silicon
- ๐ง GLM-5.3, a cheaper Gemini Flash, and NVIDIA's open router for agent pipelines
- ๐ฌ Proving agent code correct, a shell-quoting bug hiding in plain sight, and faster speculative decoding
- ๐ ๏ธ Turning your flakiest UI test into plain English
- ๐ ChatGPT ads, a new CRO at OpenAI, and Sheets learns to build dashboards
Get tomorrow's issue in your inbox.
One concise AI brief, sent after the signal clears the noise.
๐ญ THE ONE THING
โก OpenAI's Ultrafast mode pushes GPT-5.6 Sol to 14x speed, courtesy of Cerebras silicon
OpenAI opened a limited preview of "Ultrafast" on August 13, running GPT-5.6 Sol up to 14x faster than the standard API tier, topping out near 750 output tokens per second. The horsepower comes from Cerebras's wafer-scale hardware, not a new model, and Cerebras confirmed the arrangement the same day in its own investor release. Treat 14x as the ceiling, not the average: neither company has published an end-to-end number to match it, so expect real workloads to land below the headline figure. For anyone building voice agents or anything with a human waiting on the other end, that's the difference between a laggy demo and a product people trust. The limiting factor right now isn't the tech, it's access: access is capped to a small group of API customers, so plan for a waitlist before you redesign anything around it.
๐ง MODELS & RELEASES
- ๐ฅ Z.ai shipped GLM-5.3 on the same 743B-parameter base as 5.2, all the gains coming from post-training. On CyberGym it claims 84.5, ahead of Claude Mythos 5 and GPT-5.6 Sol by Z.ai's own numbers, not yet checked independently. One benchmark, not a sweep. source
- โก Google DeepMind followed 3.6 Flash with Gemini 3.7 Flash after just three weeks, pitched as the workhorse for coding and agents. Intro pricing runs $0.75/$3.75 per million tokens through year end, half the old Flash rate, with FrontierCode scores up from 34.4% to 43.6%. source
- ๐ NVIDIA paired a new 30B open MoE, Nemotron 3.5 Lightning, with NeMo Switchyard, an open-source router that spreads requests across a builder's own mix of models. NVIDIA says the combo holds frontier accuracy at close to a third of what running Opus 4.8 alone costs. Worth an independent check before you build a pipeline around a vendor's own number. source
- ๐ก Sakana AI added two models to Sakana Chat: Fugu, an orchestrator built on Gemma 4 aimed at base-model-agnostic routing, and Namazu, a refresh tuned for Japanese business tasks. Small release, real bet: that orchestration logic can outlive any one underlying model. source
- ๐ Amazon Quick landed as agentic extensions inside Word, Excel, PowerPoint, and Outlook, editing documents in place and pulling QuickSight or Salesforce data straight into them. Live now for existing Plus, Professional, and Enterprise customers, no new license required. source
๐ฌ RESEARCH HIGHLIGHTS
- Vero asks whether coding agents can do more than pass tests: can one also produce a machine-checked proof that its code matches the spec. Not "did it run," but "can it be proven correct." The gains right now come from scaffolding formal tools around the agent, not the agent getting better at proof-writing on its own. Worth tracking before agent-written code lands anywhere safety-critical. Vero
- QuoteBench caught something uglier than bad code: agents issue Bash commands through layers that serialize and reparse the model's output, and quoting can mangle a correct command after it's already been generated. A benchmark that only checks whether the task finished won't see this failure, since it happens downstream of the model entirely. If your agent's shell tool has ever done something inexplicable, this is probably why. QuoteBench
- DARTree targets a real weak spot in speculative decoding. Diffusion-based drafters guess a whole block of tokens at once for speed, but score each position independently, so the guess gets worse the further out it reaches. DARTree swaps the flat guess for a tree of autoregressive drafts, letting the verifier check more candidates per step instead of waiting on the slow base model. Inference cost coming down, if the numbers hold outside the paper. DARTree
๐ ๏ธ TRY THIS
Rewrite one brittle UI test in plain English
Faster inference (see today's lead) is what makes agent-driven testing usable in a tight loop instead of a coffee-break wait. Worth trying this week:
1. Pick the flakiest selector-based test in your suite, the one that breaks every time someone renames a div.
2. Rewrite it as a plain-English script: what a QA person would actually type in a ticket, not CSS selectors.
3. Run it against staging for a week alongside the old test, and diff the results before you trust it alone.
4. Once they agree, retire the selector version. Keep the English one as the source of truth.
Prompt: Test this checkout flow: log in as the test user, add the first item to the cart, apply promo code SAVE10, and confirm the total reflects the discount. Flag anything that looks off, not just what fails outright.Worth a look
- Amazon Nova Act First Orion swapped selector-based UI tests for plain-English ones, cutting QA cycle time and catching regressions the old scripts missed. Case study
- Bedrock AgentCore Browser Tool Drives legacy web apps through isolated browser sessions, for the systems where nobody ever built a real API. AWS walkthrough
๐ QUICK LINKS
- OpenAI started testing ads in ChatGPT's free tier, with labeling and a promise that ads won't touch the answers themselves. Testing ads in ChatGPT
- Google added a canvas mode to Sheets. Describe the dashboard or tracker you want, and it builds the layout from your data. Sheets canvas
- NVIDIA's Jensen Huang topped Glassdoor's 2026 Best CEOs list, 99% employee approval. Glassdoor ranking
- OpenAI named Dali Rajic its first Chief Revenue Officer, another sign monetization is now a standing function there, not an afterthought. Rajic appointment
- Hugging Face: 1,200+ community members ran AI agents against 2,226 ICML 2026 papers. 51% held up on at least one claim; 23% got falsified or contested. Reproduction study
See you tomorrow.
Pradeep Perugu