| 1. Claude Fable 5 (High) |
| 2. GPT 5.6 Sol (xHigh) |
| 3. Claude Opus 4.8 (Thinking) |
| 4. Kimi K3 |
| 5. Claude Sonnet 5 (High) |
The BriefAnthropic released Claude Opus 5 on Friday, a new top-of-the-line model aimed squarely at long-running agent work, and it did not raise the price: still $5 per million input tokens and $25 per million output, the same as Opus 4.8, with an optional fast mode that runs 2.5 times quicker for double the base rate. The numbers Anthropic published are about agents finishing real jobs rather than answering trivia: more than double Opus 4.8 on Frontier-Bench, three times the next-best model on ARC-AGI 3, better than Fable 5 on the OSWorld computer-use test at roughly a third of the cost, and a pass rate around 1.5 times the next model on Zapier's automation benchmark. Buried in the system card is the part that matters most if you let AI touch your files or inbox: Anthropic says this is its hardest model yet to hijack with a prompt injection. Why it matters if you work with these tools: the jump this time is in follow-through, the model checking its own work and grinding through multi-step tasks without a human nudging it every few minutes, so the honest test is not a chat window but handing it something that takes an hour.
Level UpThe cheapest win available right now is not a better model, it is spending less thinking on work that does not need it. Most current models let you dial reasoning effort up or down, and people leave it pinned high by default and pay for it on every call. Take one task you run repeatedly, run it at low, medium, and high effort, and compare the output side by side. Usually the cheap tier is fine and you keep the expensive one for the genuinely hard cases. Sebastian Raschka's write-up explains what these effort modes actually change under the hood, which makes it much easier to guess right the first time. Read: Controlling reasoning effort in LLMs (Sebastian Raschka)