| 1. Claude Fable 5 (High) |
| 2. GPT 5.6 Sol (xHigh) |
| 3. Claude Opus 4.8 (Thinking) |
| 4. Kimi K3 |
| 5. Claude Sonnet 5 (High) |
The BriefGoogle DeepMind released three Gemini models on July 21: Gemini 3.6 Flash, a faster and cheaper 3.5 Flash-Lite, and 3.5 Flash Cyber, a security-tuned variant that hunts and patches software vulnerabilities. The headline model, 3.6 Flash, uses 17% fewer output tokens than 3.5 Flash on the Artificial Analysis Index (and up to 65% fewer on one long-horizon coding benchmark) while posting higher SWE-Bench Pro and computer-use scores, priced at $1.50 per million input tokens and $7.50 output. Notably absent again: the flagship Gemini 3.5 Pro, which Google promised in June and now says it hopes to 'land soon' after missing internal performance goals, even as it confirms Gemini 4 pretraining has begun. The signal for working teams is that the cheap, fast tier is now where the real gains are landing, so the model you reach for by default is worth re-checking this quarter.
Level UpGemini 3.6 Flash cut agent token costs by up to 65% on long-horizon coding just by being more efficient, not by you changing anything. The move this week: pick your single highest-volume AI workflow (call summaries, ticket triage, drafting) and run the same 10 real inputs through both your current model and a cheaper 'Flash-class' one side by side. Compare the outputs honestly. If the cheap one holds up, you just found recurring savings with zero new engineering. See the 3.6 Flash specs, then A/B one workflow against a cheaper model