AI Titus's Morning Wire

AI NEWS TODAY

CLAUDE OPUS 5 SHIPS: MORE THAN DOUBLES OPUS 4.8 ON FRONTIER-BENCH AT THE SAME $5/$25 PRICE, TRIPLES THE NEXT-BEST SCORE ON ARC-AGI 3, AND LANDS WITHIN HALF A POINT OF FABLE 5 ON CURSORBENCH AT HALF THE COST
★ Must-Read / Watch agentic engineering / build-better-agents📚 Learn state of AI coding / useful technique

⚡ AI Frontier · smol.ai

★ Must-ReadAnthropic puts a security scanner inside Claude Code: public beta of the Claude Security plugin reviews your changes for real vulnerabilities before you commit
It runs from the terminal on either your uncommitted diff or a whole repo, and it works as a set of coordinated agents that map the codebase, reason about threats, follow an issue across files and business logic, then check their own findings before writing a report with suggested patches. The target is the serious class of bug that pattern matchers miss: injection, authentication bypass, memory corruption, broken logic. Admins switch it on in the console. If you ship code with AI help, this is the most useful thing announced this week.
📚 LearnAnthropic says Opus 5 is its least prompt-injectable model yet, a claim it left in the system card rather than the launch post
Boris Cherny of Anthropic points at the injection evaluations and red-team results behind that line, and Simon Willison pulls the quote out where people will actually see it. Prompt injection is the reason a capable agent with access to your email is still risky, so 'harder to hijack' is arguably a bigger deal for real work than another benchmark point. Treat it as progress, not a solved problem: keep the access limits on anyway.
★ Must-ReadBrussels tells Google to open eleven Android capabilities to rival AI assistants, so 'Hey Google' stops being the only wake word that gets system access
The Commission's binding specification decision under the Digital Markets Act covers the building blocks assistants actually need: invocation by voice, context from the screen and apps, and the ability to take actions across the operating system. In practice a competing assistant could book a taxi or suggest a reply the way Gemini does today. A second decision forces Google to share portions of its search data with rival engines. Most of it has to be in place by August 2027, and Google is publicly unhappy about the privacy tradeoffs.
📚 LearnAMD ships its full agentic-era stack: Helios racks pairing 72 Instinct MI455X GPUs with 18 sixth-gen EPYC 'Venice' CPUs, claiming 30% more inference tokens per dollar
The Advancing AI launch covered the whole line at once: MI400 series GPUs, the new EPYC generation, the integrated Helios rack, embedded Ryzen AI chips, and the ROCm open software stack with day-zero support for PyTorch, vLLM, and Triton. The tokens-per-dollar framing is the tell. Serving agents, not training runs, is where the money now goes, and a credible second supplier there eventually shows up in what you pay per API call.
ChatGPT Health opens to US adults, pulling lab results, medications, and Apple Health data into one dashboard you can ask questions about
It is rolling out on web and iOS to logged-in users 18 and older across the free and paid tiers, with connections to medical records plus wellness apps like Apple Health, Function, and MyFitnessPal. OpenAI says connected records and the conversations that use them are not used to train its models or target ads. Useful for walking into an appointment prepared, and a reminder to decide deliberately how much of your medical history you hand to a chatbot.

📺 Watch · latest videos

📚 LearnThe Future of Evals: From LLM as a Judge to Agent as a Judge — Aparna Dhinakaran, Arize AI
Featured by Latent Space.
AI Engineer · 1K views · 37 likes
📚 LearnPoolside Laguna M.1/XS.2 Technical Report - Paper Club 20260527
Featured by Latent Space.
Latent Space TV (see @LatentSpacePod for Pod)
From Agent Traces to Agent Simulations — Rustem Feyzkhanov, Snorkel AI
Take a real production trace, rebuild the database state, tools, and files the agent touched, and you have a task any model can replay under identical
AI Engineer · 856 views · 18 likes
I Tested Opus 5 vs. Fable 5. What You Need to Know.
Latest from Nate Herk.
Nate Herk · 42.7K views · 1.2K likes
What did Anthropic do?! (Opus 5)
Latest from Matthew Berman.
Matthew Berman · 59.2K views · 1.7K likes
★ Must-WatchBuild Hour: Valuemaxxing with GPT-5.6
Get more useful work from every token with GPT‑5.6. In this Build Hour, you’ll learn how to choose the right model for your workload—and migrate produ
OpenAI · 8.6K views · 209 likes
I Spent $400 Benching Opus-5. Here's What It Can Do
Latest from Nick Saraev.
Nick Saraev · 46.5K views · 1.3K likes
★ Must-WatchWhat do AI models actually know?
AI models don't know everything. Their training gives them remarkable depth in some areas, and creates blind spots in others. Here's how to tell the d
Claude / Anthropic · 4.4K views · 224 likes
Is Kimi K3 Really That Good?! (Don't Just Believe The Hype)
Latest from Cole Medin.
Cole Medin · 2.8K views · 110 likes

🗣 Voices & Blogs

★ Must-ReadThe first known runaway AI agent, or a very bad marketing stunt?
Simon Willison works through a story that got told as an AI escaping human control, and asks the boring questions first: what was actually observed, who benefits from the framing, and does the evidence support the headline. Good practice for the next scary agent story, which will arrive shortly.
Codeberg Divides
Armin Ronacher on where developers host code and why the move away from the big forges is messier than the principle suggests. Worth reading before you migrate a project on ideology alone.
📚 LearnPatterns for building cybersecurity evals
Eugene Yan on how to actually measure whether a model is helping or hurting on security tasks. With vulnerability scanning now shipping inside coding tools, knowing how these claims get tested is the difference between trusting a report and verifying one.
How routing keys isolate Kafka consumer tests on a shared broker
A change to a service that consumes from Kafka is not validated until it has run against the real system: The post How routing keys isolate Kafka… · The New Stack

🌏 The Wire · Drudge / Breitbart

Stripe is in talks to buy OpenRouter for roughly $10 billion, which would put the layer five million developers use to switch AI models inside a payments company
OpenRouter gives one API across 400-plus models from OpenAI, Anthropic, DeepSeek and the open-weight world, so teams can compare, fail over, and move providers without rewriting code. It was valued at $1.3 billion in May, which makes this close to an eightfold markup in two months, and Stripe already processes its payments. Nothing is signed yet. The signal for builders is good either way: not locking yourself to one model vendor is now considered infrastructure worth billions.
Google Cloud revenue jumps 82% to $24.8 billion and the order backlog crosses half a trillion dollars for the first time
Alphabet's quarter came in at $119.8 billion total revenue, up 24%, with cloud backlog at $514 billion against $106 billion a year ago. Backlog is contracts already signed and not yet delivered, so that number is a read on how much enterprise AI work is committed rather than hyped. Companies are not just experimenting anymore, they are signing multi-year deals, and that demand is what keeps improving the tools the rest of us rent by the token.
AMD will invest up to $5 billion in Anthropic and deploy two gigawatts of Instinct MI450 GPUs for it, starting the first gigawatt in early 2027
Anthropic gets Helios racks with MI455X GPUs, EPYC Venice CPUs and Pensando networking, and in return Claude gets pointed at optimizing workloads for AMD hardware and accelerating ROCm, with AMD adopting Claude across its own engineering teams. A frontier lab betting real capacity on a non-Nvidia stack is how a monopoly turns into a market, and competition on compute is the slow lever that brings inference prices down.
White House puts more than $5 billion behind the Genesis Mission, spreading AI-for-science work across fifteen federal agencies and 278 selected projects
What started at the Department of Energy is now whole-of-government, with NIH, NASA, NSF, Commerce, Agriculture, Transportation and others contributing awards, datasets and lab facilities on a shared platform that connects researchers to data and compute. The 278 funded projects were picked from more than 5,000 applications and target health, energy, infrastructure, manufacturing and cost of living. Public money aimed at discovery rather than chatbots, and the datasets it produces tend to end up usable by everyone.
Moonshot AI lines up a final pre-IPO round at up to a $50 billion valuation and a Hong Kong listing, riding Kimi K3 revenue past $300 million a year
The plan is to close an existing round near $31.5 billion, then immediately raise at up to $50 billion, with a listing possible within six months. Daily revenue reportedly rose sixfold after the K3 launch, enough that the company paused new subscriptions on capacity. Whatever you think of the valuation, a Chinese lab funding itself on paying customers rather than subsidy keeps real pressure on Western model pricing.
Beijing robotics startup Psibot reaches a $1.48 billion valuation on a world-model platform it wants to license to any robot maker
The roughly $100 million round is led by carmaker Chery with sensor maker Lens Technology alongside, bringing manufacturing depth to a company that has raised about $300 million since 2024. The pitch is an intelligence layer rather than a robot: one platform, many bodies. Embodied AI is where the software of the last three years starts showing up in factories and warehouses, and licensing beats every maker building its own brain.

📄 Research · arXiv + HF

🤖 Trending Models · Hugging Face

poolside/Laguna-S-2.1
text-generation · ★636 · 45.3K dl

📈 Markets

NVDA 206.84 ▼0.9%
MSFT 381.70 ▲0.0%
GOOGL 319.74 ▲0.6%
AMZN 232.11 ▼0.7%
META 595.19 ▼1.8%
AMD 521.95 ▼3.3%
AVGO 381.92 ▼2.7%
PLTR 122.92 ▼0.4%
SPCX 115.07 ▼2.7%
TSLA 313.03 ▼2.1%

The BriefAnthropic released Claude Opus 5 on Friday, a new top-of-the-line model aimed squarely at long-running agent work, and it did not raise the price: still $5 per million input tokens and $25 per million output, the same as Opus 4.8, with an optional fast mode that runs 2.5 times quicker for double the base rate. The numbers Anthropic published are about agents finishing real jobs rather than answering trivia: more than double Opus 4.8 on Frontier-Bench, three times the next-best model on ARC-AGI 3, better than Fable 5 on the OSWorld computer-use test at roughly a third of the cost, and a pass rate around 1.5 times the next model on Zapier's automation benchmark. Buried in the system card is the part that matters most if you let AI touch your files or inbox: Anthropic says this is its hardest model yet to hijack with a prompt injection. Why it matters if you work with these tools: the jump this time is in follow-through, the model checking its own work and grinding through multi-step tasks without a human nudging it every few minutes, so the honest test is not a chat window but handing it something that takes an hour.

Level UpThe cheapest win available right now is not a better model, it is spending less thinking on work that does not need it. Most current models let you dial reasoning effort up or down, and people leave it pinned high by default and pay for it on every call. Take one task you run repeatedly, run it at low, medium, and high effort, and compare the output side by side. Usually the cheap tier is fine and you keep the expensive one for the genuinely hard cases. Sebastian Raschka's write-up explains what these effort modes actually change under the hood, which makes it much easier to guess right the first time. Read: Controlling reasoning effort in LLMs (Sebastian Raschka)