| 1. Claude Fable 5 (High) |
| 2. Claude Opus 4.8 (Thinking) |
| 3. GPT 5.5 (xHigh) |
| 4. Claude Opus 4.7 (Thinking) |
| 5. Claude Opus 4.7 |
The BriefOpenAI released GPT-Live, a new voice model built to be 'full duplex,' meaning it can listen and speak at the same time instead of taking turns like a walkie-talkie. That lets you interrupt it mid-sentence and have it react naturally, closer to how a real phone call works. The smaller GPT-Live-1 mini is now the default voice mode for everyone in ChatGPT, while the bigger GPT-Live-1 goes to paid subscribers, and it matters because voice is quietly becoming a primary way people use AI, not just typing.
Level UpThis week's headlines are full of AI-agent security incidents (see The Wire below), so the skill to build now is testing whether your own agents can be tricked. Eugene Yan lays out practical patterns for cybersecurity evals: how to probe an agent for prompt injection, data leaks, and misuse before it ships. If you run any agent that touches customer data or calls tools, spend an hour writing a few of these adversarial tests this week. Patterns for Building Cybersecurity Evals (Eugene Yan)