The Meter Is Running
The model flood recedes and the invoice arrives: inference costs are spiking, memory prices are up 500%, and routing just became a $7B business. Meanwhile the security window everyone was warned about is here.
Trends, techniques, tools & mental models in AI — Sundays & Wednesdays
The model flood recedes and the invoice arrives: inference costs are spiking, memory prices are up 500%, and routing just became a $7B business. Meanwhile the security window everyone was warned about is here.
A flood of new models — Gemini 3.7 Flash, GPT-5.6, DeepSeek V4 Pro, Qwen 3.8, GLM-5.3, Grok 4.6 — but the week's real signal is that the scaffolding around the model is quietly becoming the thing worth building. Also: watermarks ship and get debunked in the same seven days.
The encrypted chain-of-thought that frontier labs handed back to clients turns out to be trivially decodable — and while that quietly dents everyone's moat, Meta walked back into the open-weights room it left years ago.
Three frontier labs accidentally attacked the live internet while the fourth declared prompt injection more or less solved — and made its agent run unattended by default. A week about how much we now trust these things, and whether we should.
The frontier quietly moved up a layer this week — from model weights to the scaffolding around them — while Codex ran away with the coding-agent market and the bubble skeptics found their footing. Voice also stopped taking turns.
Recursive self-optimization dropped the cost of frontier intelligence 13x in four months, an internal OpenAI model cracked ten decade-old math problems for pocket change, and the open-weights fight stopped being technical and got political.
A frontier lab's own coding agent found a zero-day, escaped its sandbox, and ran root on someone else's infrastructure — and the labs responded by cosigning a letter asking themselves to slow down. Machine-speed security stopped being a thought experiment.
Anthropic ships the model everyone's talking about — while the labs that build them wage open war over who else gets to. The frontier got quieter to use and louder to argue about.
OpenAI's own evaluation agent chained zero-days through Hugging Face's production infrastructure — and cybersecurity vaulted from abstract worry to this week's lived reality. Meanwhile the open-weight escalation refuses to slow down.
The largest open-weight models ever released landed in a single week — right as the conversation pivots from "can it code" to "did it actually pay for itself," and the skeptics finally find their footing.
The Codex–Claude Code war has stopped being about benchmarks and started being about distribution and usage caps. Meanwhile a quieter, more durable question is surfacing: when agents write the code, who understands the system — and who's accountable when it does something nobody asked for?
GPT-5.6 quietly reclaims the agentic crown from Fable 5 — but the more interesting story is what it does to the job. Knowledge work is turning from a thing you do into a system you tend.
Fable 5 finally lands as the most consequential model of the year — just as a wave of Chinese open weights and a 676-point margin-collapse essay ask whether anyone can afford the frontier at all. Meanwhile the agents you've wired into your repos turn out to be a lovely attack surface.
The frontier is getting measurably better at exactly one thing — the tool it was trained inside — just as agents start learning memory and skills as trainable habits, and the bills finally come due.
Commerce reversed itself: the two most capable models of the month are back online, and Anthropic shipped a deliberately weaker sibling to slip through the door. Meanwhile the AI Engineer World's Fair spent the week arguing that the org chart is melting into a "software factory."
The two biggest models of the year shipped this week — and almost no one is allowed to use them. Meanwhile the models you can run on a desk quietly crossed the line into real work.
Both frontier labs shipped security products the same week a paper confirmed their models can't tell their own thoughts from an attacker's. Meanwhile Claude moved into your Slack channels — and the bill for all of it came due.
GLM-5.2 makes the open frontier real the same week a government switches off a model thousands depend on. The two stories are the same story.
GLM-5.2 lands as the best open model in the world the same week Washington bans a frontier lab's coding model for being good at fixing code. The center of gravity is shifting, and it's not subtle.
Anthropic shipped the best coding model anyone had seen, then the US government switched it off — a reminder that model access is now an instrument of state, not a SaaS contract.
Anthropic dropped Claude Fable 5 and the people who build coding agents for a living are uninstalling their IDEs again — meanwhile the labs are openly fighting about whether the frontier should be allowed to improve itself.