The Price of Thinking Halves Again
A new wave of models cut prices 40-50% in a single week, Anthropic clawed back its defectors with Opus 5.5 — and a genuinely new shape of model, the “decision model,” started to look like more than a curiosity.
The Big Picture
The headline is arithmetic. In the space of a few days we got Grok 4.7, MiMo v2.6, Claude Opus 5.5, and OpenAI’s GPT-6 Sol and Luna — and nearly everyone cut prices 40-50%. Simon Willison’s pricing table tells the story: GPT-6 Luna lands at $0.10/M input and $0.50/M output — half its already-cheap predecessor — while GPT-6 Sol delivers most of Astra’s power at roughly a fifth of the cost. The r/LocalLLaMA crowd is passing around an Epoch chart showing the cost of intelligence falling ~50% per quarter. If that curve holds, the planning horizon for anything you build against a fixed model price is now about a season long.
The more interesting story is competitive, not numeric. Claude Opus 5.5 is the surprise of the cycle — it fixed the “hard to love” personality problems of Opus 5 and is pulling Codex converts back to Claude. Meanwhile GPT-6 Sol has become the quiet daily driver for people who care about speed and cost. The takeaway for how you work: the “best model” question has fully collapsed into “best model for this task, at this price,” and the only way to answer it is to measure on your own work — which is exactly why personal benchmarks are having a moment (below).
Underneath the model churn, a second, structural shift is surfacing: as generation gets cheap and fast, the bottleneck moves everywhere else — CI pipelines, code review, human attention. And a new class of model, the “decision model,” is being pitched as the tool for the parts of the stack where you never actually wanted prose in the first place.
Themes
The decision model arrives — a genuinely new shape
TypeSafe AI’s Jev is the most novel thing this period, and worth understanding even if you never use it. It takes text in but returns floating-point numbers — yes/no answers, category choices, scores, each with a confidence — rather than tokens. Willison, following Maggie Appleton, prefers “decision model” to TypeSafe’s “System One model.” The economics are the hook: input priced at $0.042/M, output free, because there barely is any. Think “frontier-intelligence function call: unstructured state in, typed probabilistic decisions out.”
The early evidence suggests this isn’t just a pricing gimmick. A JEV-as-a-Judge paper found a decision-only judge lands within three points of a state-of-the-art LLM judge at 0.36% of the cost, with a cascade that escalates only low-confidence calls retaining 99% of accuracy. The pattern to watch: use a cheap decision model as a confident first pass, reserve the expensive generative model for the hard forks. The community reaction is its own signal — Jev in 25 Lines of Python hit 508 points on HN, and r/LocalLLaMA is openly complaining about Jev shill posts, which is what a real launch and an aggressive marketing budget look like at the same time. Confidence: clearly landing as a category; verdict still out on how much of your classification/routing work it can actually absorb.
Go deeper: Willison’s breakdown · the definitive Jev podcast · JEV-as-a-Judge · llm-typesafe plugin
Benchmarks are personal now
With four credible daily-driver models at wildly different prices, generic leaderboards have never been less useful. Mike Taylor’s guide to building your own benchmark makes the case bluntly — quoting Mollick, MMLU-Pro asks models the cranial capacity of Homo erectus, which tells you nothing about whether a model can do your job. The practice that’s emerging: keep a private eval set drawn from your actual tasks, and re-run it on each release. Lenny Rachitsky’s blind taste test across writing, frontend, agents, and SVGs is the same instinct applied publicly.
The research frontier is moving from “did the agent succeed” to “did it make good decisions along the way.” Taste-Bench (111 upvotes) measures an agent’s taste — its ability to pick the better branch at a decision fork without seeing the outcome — and finds the best frontier model gets only 59.7% right. That’s a useful reframe: as agents run longer, the quality of intermediate choices dominates, and it’s a distinct capability from raw benchmark score.
Go deeper: Build a personal benchmark · Advanced evals: finding hidden failures · Taste-Bench
When generation is free, everything else is the bottleneck
The clearest second-order effect of cheap, fast code generation: the constraints move downstream. Linear’s engineering post, AI coding has made CI a bottleneck (312 points), is the canonical example — when agents produce PRs faster than your pipeline can test them, the pipeline becomes the thing you re-architect. Warp’s counterpoint is to lean all the way in: 2,000 PRs a month via “AI software factories” that turn Slack, Linear, and GitHub requests into tracked, tested PRs and measure the output over time.
There’s a human cost surfacing in parallel, and it’s not pretty. Michael Heap’s I don’t want the details (212 points) and the widely-shared voxium quote — “people are working 12 to 13 hours a day just to press enter… nobody is reading anything” — describe the same failure mode: throughput treated as the goal, comprehension quietly abandoned. The rebuttal, if AI coding is lowering your code quality, you’re not managing quality right, is worth reading as the optimistic counter. This is the labor reckoning from last week, now showing up as a concrete engineering-management problem rather than a vibe.
Go deeper: CI as the new bottleneck · Warp’s AI factories · I don’t want the details
Agents get wallets, and the web starts pushing back
The “agent as a teammate with its own accounts” idea is now shipping products: Solid gives agents their own computers, accounts, and budgets; Plane Agents lets you assign work to agents like any team member; and Strands Harness is a new runtime for building them. But the friction is arriving on schedule — Amazon blocked Meta’s Muse agent from shopping on its site, an early skirmish in what will be a long fight over who gets to let autonomous buyers onto their platform. Meanwhile Willison offered a sharp defense of MCP against the “always a bad idea” crowd: full YOLO terminal agents don’t need it, but anything with access controls, auth that hides keys, and audit logging still wants exactly what MCP provides. Early signal, but the “agents transacting on the open web” story is about to become a governance story.
Go deeper: MCP defense · Amazon blocks Muse · Strands Harness
Radar
- Claude Opus 5.5 — Anthropic’s fix for Opus 5’s personality problems; matching Fable 5.1 on many tasks while cheaper, and winning back Codex defectors.
- GPT-6 Sol & Luna — Two new OpenAI models at roughly half the price of their GPT-5.6 equivalents; Luna at $0.10/M input is the new cheap-app default.
- Jev — TypeSafe’s “decision model” returning typed probabilistic outputs at $0.042/M input, output free; 508 points on HN for a 25-line implementation.
- MiMo-V2.6-Pro — Xiaomi’s 1T-A42B open-weights model trained for ~$3M, topping open leaderboards — though senior devs call it benchmaxxed on real tasks. Verify before you trust the numbers.
- HySparse2 / MiMo-V3 — The sparse-attention core of Xiaomi’s next architecture, out now as a paper; worth a look if you track efficiency techniques.
- Transformers now runs GGUF quants natively — Load llama.cpp quants directly with standard Transformers APIs; smaller footprint plus PyTorch debugging tooling.
- Cloudflare Python Workers (GA) — Python via Pyodide/WebAssembly is now first-class on Cloudflare’s edge, with a full local simulation dev environment.
- univer — “The Office Harness for AI Agents” — spreadsheets, docs, slides, and PDF in one runtime; 768 stars this week and a clean fit for document-manipulating agents.
- Gemini 3.8 TTS — Google’s new text-to-speech, part of a broad voice-and-translation wave (see also NVIDIA’s Nemotron diarization for real-time multi-speaker).
- Stripe’s Knowledge AI Platform — How Stripe turns internal knowledge into agent-usable context (118 points); a useful reference architecture for institutional memory.
- GPT-6 Astra learns to drive — A driving benchmark for the top-end model; more a capability curiosity than a product, but a notable data point on multimodal control (175 points).
- Taste-Bench — Measures the intermediate decision quality (“taste”) of long-horizon agents; best frontier model scores 59.7%.
Don’t Miss
- The Pentagon blames AI overreliance for a missile strike on an Iran school (849 points). An update to last week’s “receipts” theme, and the most serious one yet: the first high-profile military incident where the after-action report names overreliance on AI as a contributing cause. Pair it with the FT finding that chatbots give wrong answers to financial queries most of the time — the failure modes are graduating from embarrassing to consequential.
- AI Has No Wisdom and Neither Will You (383 points, 539 comments). The mental-model essay of the week — a sharp argument that offloading judgment atrophies the very capacity you’re offloading. Read it alongside the “I don’t want the details” thread.
- The current balance of power in open models. Nathan Lambert’s Congressional testimony, expanded — the clearest current read on the US-China open-weights gap, timely given MiMo’s ascent.