| 1. Claude Opus 5 (High) |
| 2. Claude Fable 5 (High) |
| 3. Claude Opus 5 (Max) |
| 4. GPT 5.6 Sol (xHigh) |
| 5. Kimi K3 (Max) |
The BriefAnthropic retuned the filter that decides when a question to Claude touches biology, and says false alarms dropped about 85 percent. In plain terms, far fewer people asking something ordinary, like what a lab result means, get quietly bounced to a weaker backup model partway through the conversation. The genuinely dangerous requests are still blocked. What changed is the filter's ability to tell a worried patient apart from someone trying to cause harm, which is the part these systems have been worst at.
Level UpTest whether a small cheap model can do one of your expensive jobs. A team at Castform took a four billion parameter open model, trained it on synthetic questions generated from their own product docs, and got retrieval accuracy matching a frontier model at roughly one hundredth the cost. Pick your narrowest, highest-volume task and try it there first. The write-up shows the actual training setup, so this is a weekend project and not a research program. How Castform beat frontier models on price and efficiency