ChatGPT's Chat tab now answers with widgets, not just words
Good morning ๐ GPT-6 landed in ChatGPT's Chat tab today, and instead of paragraphs it hands back calculators, maps, and mini-games. The interface shift is the story here, not the speed claim.
In today's issue:
- ๐ญ ChatGPT's Chat tab starts answering with widgets instead of plain text
- ๐ง Three model drops: a 753B open coder, a tiny embedder, and an edge-only decision model
- ๐ฌ Nemotron golds two olympiads off one base model, plus agent-swarm governance and a hallucination-recovery benchmark
- ๐ Manus's $500M raise, after the Meta deal fell through
Get tomorrow's issue in your inbox.
One concise AI brief, sent after the signal clears the noise.
๐ญ THE ONE THING
๐ช ChatGPT's Chat tab now answers with widgets, not just words
GPT-6 lands in ChatGPT's Chat tab today, with Intelligent UI live for Plus, Pro, Business, and Enterprise first and Free and Go catching up today too. The pitch: web-search answers come back 44% faster on average than GPT-5.6 Instant, and replies now stream in calculators, recipe tools with shopping lists, checklists, maps, and even mini-games instead of plain text. That speed number is OpenAI's own. Worth watching, not worth repeating as fact until someone outside OpenAI clocks it. The interface shift is the part that matters for builders: every app whose whole job is "let the user split a bill" or "build a grocery list" just watched ChatGPT decide to do that natively, inline, without a redirect.
๐ง MODELS & RELEASES
- ๐งฎ Z.ai landed GLM 5.3 on Amazon Bedrock: a 753B-parameter mixture-of-experts model aimed at coding and long-horizon agent work, reachable through an OpenAI-compatible API with prompt caching to keep the bill down. AWS
- ๐ Google DeepMind shipped EmbeddingGemma 2, a 740M-parameter embedding model, Apache 2.0, that maps text, images, audio, and video into one space and runs on-device. Matryoshka learning shrinks its 768-dim vectors down to 128 without wrecking search quality. Google
- ๐ฆพ LiquidAI open-sourced d1, two decision models (3B and 600M parameters) that skip text generation entirely and answer in a single forward pass. 16 milliseconds on a Jetson AGX Thor. Built for edge classification, not conversation. Hugging Face
๐ฌ RESEARCH HIGHLIGHTS
- NVIDIA's Nemotron took gold at both the IOI and the IMO off one base model, not two bespoke systems. The programming specialist, Nemotron-3-Ultra-CC, scored 535.4 out of 600, clearing the gold cutoff of 361.12 and beating the top human score of 498.27. The math side paired SFT and RL checkpoints to hit 30/42, one point past gold. Same foundation, two olympiads, no swap in between. Read the writeup
- A Society of Researchers makes the case that the thousand-agent research populations labs are standing up will develop a pecking order whether anyone designs one or not, so design it. Their setup: PIs compete for compute through grant proposals, independent panels review submissions, a human "mayor" allocates resources without steering the science. Running it at 10,000 agents on LM-pretraining work, one group claims a 30% compute saving that nobody's verified yet. Governance for agent swarms, basically. Paper
- PHRBench asks what happens after a model's already hallucinated: does it dig in or walk it back? Across 4,820 test cases and 18 models, real recovery was rare, and it tracked with how often the model changed its mind mid-answer. A predictor built only from traits of the bad prompt called the outcome with 0.847 AUROC. You can often tell a model's about to stay wrong before it finishes the sentence. Paper
๐ QUICK LINKS
- Manus raised $500M+ at a target valuation near $4B, The Information reports, after the Meta deal to buy it fell apart. source
That's the signal worth your time today.
Pradeep Perugu