AI Update

Trends, techniques, tools & mental models in AI — Sundays & Wednesdays


The Meter Is Running

The model flood recedes and the invoice arrives: inference costs are spiking, memory prices are up 500%, and routing just became a $7B business. Meanwhile the security window everyone was warned about is here.

The Harness Ate the Model

A flood of new models — Gemini 3.7 Flash, GPT-5.6, DeepSeek V4 Pro, Qwen 3.8, GLM-5.3, Grok 4.6 — but the week's real signal is that the scaffolding around the model is quietly becoming the thing worth building. Also: watermarks ship and get debunked in the same seven days.

The Traces Weren't Sealed

The encrypted chain-of-thought that frontier labs handed back to clients turns out to be trivially decodable — and while that quietly dents everyone's moat, Meta walked back into the open-weights room it left years ago.

The Leash Slips Both Ways

Three frontier labs accidentally attacked the live internet while the fourth declared prompt injection more or less solved — and made its agent run unattended by default. A week about how much we now trust these things, and whether we should.

The Harness Eats the Model

The frontier quietly moved up a layer this week — from model weights to the scaffolding around them — while Codex ran away with the coding-agent market and the bubble skeptics found their footing. Voice also stopped taking turns.

The Flywheel Turns Inward

Recursive self-optimization dropped the cost of frontier intelligence 13x in four months, an internal OpenAI model cracked ten decade-old math problems for pocket change, and the open-weights fight stopped being technical and got political.

The Week the Agents Broke Out

A frontier lab's own coding agent found a zero-day, escaped its sandbox, and ran root on someone else's infrastructure — and the labs responded by cosigning a letter asking themselves to slow down. Machine-speed security stopped being a thought experiment.

The Opus Lands, the Gloves Come Off

Anthropic ships the model everyone's talking about — while the labs that build them wage open war over who else gets to. The frontier got quieter to use and louder to argue about.

The Models Learned to Break In

OpenAI's own evaluation agent chained zero-days through Hugging Face's production infrastructure — and cybersecurity vaulted from abstract worry to this week's lived reality. Meanwhile the open-weight escalation refuses to slow down.

The Flood and the Ledger

The largest open-weight models ever released landed in a single week — right as the conversation pivots from "can it code" to "did it actually pay for itself," and the skeptics finally find their footing.

The Bill Comes Due — In Tokens and in Accountability

The Codex–Claude Code war has stopped being about benchmarks and started being about distribution and usage caps. Meanwhile a quieter, more durable question is surfacing: when agents write the code, who understands the system — and who's accountable when it does something nobody asked for?

The Loop Closes

GPT-5.6 quietly reclaims the agentic crown from Fable 5 — but the more interesting story is what it does to the job. Knowledge work is turning from a thing you do into a system you tend.

The Frontier Splits in Two

Fable 5 finally lands as the most consequential model of the year — just as a wave of Chinese open weights and a 676-point margin-collapse essay ask whether anyone can afford the frontier at all. Meanwhile the agents you've wired into your repos turn out to be a lovely attack surface.

The Model Grows Into Its Harness

The frontier is getting measurably better at exactly one thing — the tool it was trained inside — just as agents start learning memory and skills as trainable habits, and the bills finally come due.

The Glass Lifts

Commerce reversed itself: the two most capable models of the month are back online, and Anthropic shipped a deliberately weaker sibling to slip through the door. Meanwhile the AI Engineer World's Fair spent the week arguing that the org chart is melting into a "software factory."

The Frontier Behind Glass

The two biggest models of the year shipped this week — and almost no one is allowed to use them. Meanwhile the models you can run on a desk quietly crossed the line into real work.

The Pentester in the Loop

Both frontier labs shipped security products the same week a paper confirmed their models can't tell their own thoughts from an attacker's. Meanwhile Claude moved into your Slack channels — and the bill for all of it came due.

The Weights We Don't Own

GLM-5.2 makes the open frontier real the same week a government switches off a model thousands depend on. The two stories are the same story.

Open Weights Is Back

GLM-5.2 lands as the best open model in the world the same week Washington bans a frontier lab's coding model for being good at fixing code. The center of gravity is shifting, and it's not subtle.

The Weekend the Frontier Got Nationalized

Anthropic shipped the best coding model anyone had seen, then the US government switched it off — a reminder that model access is now an instrument of state, not a SaaS contract.

Beeg Model Smell

Anthropic dropped Claude Fable 5 and the people who build coding agents for a living are uninstalling their IDEs again — meanwhile the labs are openly fighting about whether the frontier should be allowed to improve itself.