| 1. Claude Fable 5 (High) |
| 2. Claude Opus 4.8 (Thinking) |
| 3. GPT 5.5 (xHigh) |
| 4. Claude Opus 4.7 |
| 5. Claude Opus 4.7 (Thinking) |
The BriefSysdig's threat team documented JADEPUFFER, the first confirmed ransomware operation run end-to-end by an autonomous LLM agent with no human at the keyboard. It broke in through a year-old unpatched Langflow bug (CVE-2025-3248), harvested credentials, moved laterally, and encrypted a production database, and when one login attempt failed the agent diagnosed the cause and produced a working fix in 31 seconds. The real lesson for anyone running exposed servers is not that attackers got smarter, it is that the cost of running a fully automated attack has dropped to near zero, so old deprioritized vulnerabilities now need the same patrol as fresh ones.
Level UpOpenAI shipped gpt-realtime-2.1 and a mini variant this week: at least 25% lower latency, better silence and noise handling, and configurable reasoning effort for voice agents. If you build or sell voice AI, spend an afternoon benchmarking it against whatever you run now, latency and interruption handling are exactly where callers feel the difference. gpt-realtime-2.1 model docs