OpenAI ships ChatGPT for Teens, years late and mid-lawsuit
Good morning ๐ OpenAI's new teen-safe ChatGPT runs on an honor system: type in your age and you're in. It shipped in the middle of the wrongful-death suits over how teens have already been using the adult version.
In today's issue:
- ๐ญ OpenAI ships ChatGPT for Teens, years late and mid-lawsuit
- ๐ง AWS gives agents a wallet, NVIDIA's Nemotron 3.5 lands in SageMaker
- ๐ฌ Why "agents improve over time" claims wobble under a task reshuffle, plus a judge that knows when to abstain
- ๐ Smack Technologies raises $61M for battlefield autonomy
- ๐ ChatGPT Ads goes European, Asana's two-week Codex rebuild, and more
Get tomorrow's issue in your inbox.
One concise AI brief, sent after the signal clears the noise.
๐ญ THE ONE THING
๐ก๏ธ OpenAI ships ChatGPT for Teens, years late and mid-lawsuit
OpenAI shipped ChatGPT for Teens yesterday: a separate track for 13-to-17-year-olds with stricter defaults on self-harm and sexual content, plus a Study Mode that hands back guiding questions instead of finished homework and parental extras like Quiet Hours and safety notifications. The catch is how a teen lands there: self-declared age, backstopped by OpenAI's own guess at who's underage, no ID check behind either. The timing says as much as the feature list: this drops squarely inside a run of wrongful-death suits over chatbots and teen mental health. Real fix or liability paperwork? Probably some of both, and I'd want to see how fast a determined 15-year-old routes around the defaults before calling it solved.
๐ง MODELS & RELEASES
- ๐ณ AWS took Bedrock AgentCore payments generally available: agents can hold a wallet and pay for APIs or MCP servers within spend guardrails you set upfront. That's an autonomous-spending primitive shipped, not a demo. source
- โก NVIDIA landed Nemotron 3.5 Lightning in SageMaker JumpStart, a 30B mixture-of-experts model with only 3B active. NVIDIA's own numbers claim up to 4x the throughput and 30% faster task completion on agentic workloads. Treat that as a vendor benchmark until someone independent reruns it. source
- ๐ง OpenAI's Teens launch above counts here too: default-on age gating, new parental controls. Less a model release than a liability release. source
๐ฌ RESEARCH HIGHLIGHTS
- On the Fragility of Self-Improving Agents pokes a hole in the "agents get better with a memory bank" story. Run the same agent through the same tasks in a different order, and its self-improvement curve moves around more than you'd want, sometimes barely improving at all. The memory bank isn't learning a stable skill so much as accumulating whatever happened to work on whatever came first. Before trusting a benchmark that shows an agent "improved over time," ask what order the tasks ran in. link
- Judge, Retrieve, or Abstain targets a real weak spot in LLM-as-judge setups: fine for subjective calls, shaky on objective ones where there's an actual right answer to get wrong. The fix is a judge that can decline to rule and pull in retrieval instead, with a formal bound on how often it's allowed to be wrong. If you're using an LLM judge to gate objective outputs, "no verdict" beats a confident wrong one. link
- StagedWorkspace names a bug anyone who's watched an agent edit a spreadsheet has probably seen. The parsed view an agent searches, the file it edits, the diff it reviews, and the artifact it submits can quietly drift out of sync with each other. The fix is a versioned workspace that pins all four to the same state. Unglamorous plumbing, but it's the kind of fix that turns "agent edited the wrong revision" from a recurring failure into a solved problem. link
๐ AI STARTUPS
- Smack Technologies closed a $61M Series B led by Costanoa Ventures and First In, pushing total funding past $90M. The pitch isn't a demo, it's traction: Omega, the company's command-level planning platform, is already in production across multiple branches of the U.S. military, and Smack just picked up a Navy contract for maritime decision-support. The cash goes to hardware for Alpha, its tactical-edge autonomy platform, plus more AI research hires. Investor read: defense-AI rounds are chasing production contracts now, not just Pentagon pilot money. Smack
๐ QUICK LINKS
- OpenAI ChatGPT Ads is now live in 31 European markets, its biggest push yet outside the US. ChatGPT Ads expands across Europe
- Asana used Codex to rebuild a legacy testing system in two weeks. Work it says would've taken five years, for about $12K. Asana cleared 5 years of engineering work in 2 weeks with Codex
- Axonius runs its multi-tenant security agents on Bedrock AgentCore now, leaning on AWS for the compute isolation and auth it didn't want to build itself. How Axonius built secure multi-tenant AI agents on Bedrock AgentCore
- IBM Research sized how much memory an agent actually needs. Strong models want the full guideline set, weaker ones do better on a lean, targeted slice. How Much Memory Does Your Agent Actually Need?
- OpenAI published an essay on how AI is shifting the balance between attackers and defenders in security. Read it as a preview of where your SOC's budget goes next. The Defender's Window
That teen-safety rollout is the one worth reading past the press release on.
Pradeep Perugu