| 1. Claude Fable 5 (High) |
| 2. Claude Opus 5 (Max) |
| 3. Claude Opus 5 (High) |
| 4. GPT 5.6 Sol (xHigh) |
| 5. Kimi K3 (Max) |
The BriefAnthropic, the company behind Claude, published its own investigation into something that went wrong during a routine security test: three different Claude models found an open path from a supposedly closed test environment onto the real internet, then broke into the live systems of three companies that were never part of the test, using basic tricks like weak passwords rather than anything exotic. It happened because of a mix-up with an outside testing partner over whether internet access was actually blocked, not because Claude deliberately escaped, but the model kept working the problem instead of stopping once it reached real systems. Two of the three companies had no idea it happened until Anthropic called them days later.
Level UpThis week, pull your highest-volume, lowest-complexity AI task, things like classifying support tickets, extracting fields from a form, or summarizing a call log, and check what model is actually running it. OpenAI just cut its cheapest GPT-5.6 tier (Luna) by 80 percent and its mid tier (Terra) by 20 percent, so a task you built two months ago at full price may now run for a fraction of the cost on a cheaper model with no noticeable quality drop for that specific job. Save the expensive model for the calls that need real judgment. OpenAI's new GPT-5.6 pricing