AI Titus's Morning Wire

AI NEWS TODAY

GOOGLE SHIPS THREE GEMINI MODELS AT ONCE: 3.6 FLASH, 3.5 FLASH-LITE, AND A SECURITY-TUNED FLASH CYBER, WHILE 3.5 PRO STAYS DELAYED
★ Must-Read / Watch agentic engineering / build-better-agents📚 Learn state of AI coding / useful technique

⚡ AI Frontier · smol.ai

★ Must-ReadGemini 3.6 Flash does more work with fewer tokens: 17% leaner output, higher coding and computer-use scores, same price ceiling
The new workhorse posts 58.7% on SWE-Bench Pro (up from 55.1%) and 83.0% on OSWorld-Verified computer-use (up from 78.4%), while using 17% fewer output tokens than 3.5 Flash per the Artificial Analysis Index. Priced at $1.50 per million input tokens and $7.50 output. For anyone running agents at volume, the efficiency gain is a direct cost cut without a quality tradeoff.
📚 LearnGemini 3.5 Flash Cyber is a vulnerability hunter that finds, validates, and patches code exploits at scale
Built on 3.5 Flash and fine-tuned for security, it powers Google's automated CodeMender platform, which calls it repeatedly to analyze code and catch exploits before attackers do. It ships first to governments and trusted partners through a limited-access pilot. A clear signal that defensive security is becoming a first-class model specialty, not a prompt you bolt on.
Nvidia publishes Vera CPU details and SPEC CPU 2026 numbers, aiming its custom core at agentic workloads
Vera packs 88 custom Olympus cores tuned for single-thread performance and memory bandwidth rather than raw core count, and edged a dual-socket AMD Epyc 9755 on SPECrate integer (925 vs 898) with fewer threads. Nvidia notes the run is unofficial since Vera is not broadly available yet. The design bet: agent workloads care more about fast single threads and bandwidth than core density.
Google's flagship Gemini 3.5 Pro slips again as Gemini 4 pretraining begins
Promised for June, 3.5 Pro is still testing with partners after missing internal performance targets; product lead Logan Kilpatrick says it hopes to 'land soon' and confirms the team has started its 'most ambitious pre-training run yet' for Gemini 4. The pattern across labs: shipping the fast, cheap tier on schedule while the top-end flagship keeps slipping.

📺 Watch · latest videos

How I’d Make Money with Claude if my life depended on it
Latest from Nate Herk.
Nate Herk · 1.8K views · 136 likes
Kimi K3 Designs Websites That Feel Like MOVIES For Just $1
Latest from Nick Saraev.
Nick Saraev · 23.6K views · 1.1K likes
The Most Important Conversation in AI Right Now
Latest from Matthew Berman.
Matthew Berman · 93.7K views · 3.1K likes
★ Must-WatchWhat is sycophancy?
A model that always agrees with you isn’t a helpful model. See why sycophancy matters, how you can spot it when chatting with AI, and what we’re doing
Claude / Anthropic · 6.1K views · 430 likes
★ Must-WatchPlugins in ChatGPT
Plugins connect ChatGPT to the tools you use for email, cloud storage, calendars, work chat, and more. See how to browse the plugin directory, choose
OpenAI · 18.7K views · 614 likes
Engineers... STOP Picking GPT-5.6 Sol OR Claude Fable 5… FUSE THEM
The AI industry wants you to pick a winner: GPT 5.6 Sol OR Claude Fable 5. 🔥 That's the trap. Picking one model is the single biggest mistake you can
IndyDevDan · 27.3K views · 1.1K likes
Tmux + Fable = Cut 35% less token
Open Agent Teams skill + my Claude.md setup: https://github.com/AI-Builder-Club/skills/blob/main/skills/open-agent-teams/SKILL.md 🔗 Links - 1-hr deep
AI Jason · 6.9K views · 196 likes

🗣 Voices & Blogs

Moonshot launched Kimi K3. Then demand shut down subscriptions in 48 hours.
Moonshot AI became the latest AI company to discover that launching a popular model is only half the battle. Less The post Moonshot launched Kimi K3… · The New Stack

🌏 The Wire · Drudge / Breitbart

Microsoft commits multibillion dollars to expand Mistral in Europe on thousands of Nvidia Vera Rubin GPUs
The expanded partnership brings Mistral Medium 3.5 and the OCR 4 document model into Microsoft Foundry and Copilot Studio, and lets regulated European enterprises run them across cloud, cloud-connected, and fully disconnected Azure Local environments. It is a sovereign-AI play: frontier capability customers can keep inside their own walls. For MSPs and regulated clients, on-prem and air-gapped AI just got a major backer.

🤖 Trending Models · Hugging Face

poolside/Laguna-S-2.1
text-generation · ★313 · 3.1K dl

📈 Markets

NVDA 207.29 ▲2.0%
MSFT 397.75 ▼1.1%
GOOGL 347.15 ▼1.4%
AMZN 247.55 ▼1.0%
META 643.81 ▼0.3%
AMD 544.43 ▲8.1%
AVGO 386.50 ▲2.2%
PLTR 132.66 ▼1.6%
SPCX 123.54 ▲3.1%
TSLA 378.93 ▲2.5%

The BriefGoogle DeepMind released three Gemini models on July 21: Gemini 3.6 Flash, a faster and cheaper 3.5 Flash-Lite, and 3.5 Flash Cyber, a security-tuned variant that hunts and patches software vulnerabilities. The headline model, 3.6 Flash, uses 17% fewer output tokens than 3.5 Flash on the Artificial Analysis Index (and up to 65% fewer on one long-horizon coding benchmark) while posting higher SWE-Bench Pro and computer-use scores, priced at $1.50 per million input tokens and $7.50 output. Notably absent again: the flagship Gemini 3.5 Pro, which Google promised in June and now says it hopes to 'land soon' after missing internal performance goals, even as it confirms Gemini 4 pretraining has begun. The signal for working teams is that the cheap, fast tier is now where the real gains are landing, so the model you reach for by default is worth re-checking this quarter.

Level UpGemini 3.6 Flash cut agent token costs by up to 65% on long-horizon coding just by being more efficient, not by you changing anything. The move this week: pick your single highest-volume AI workflow (call summaries, ticket triage, drafting) and run the same 10 real inputs through both your current model and a cheaper 'Flash-class' one side by side. Compare the outputs honestly. If the cheap one holds up, you just found recurring savings with zero new engineering. See the 3.6 Flash specs, then A/B one workflow against a cheaper model