The Fast Half of the Brain
A new model category — one that decides instead of talks — became the week’s biggest story, while voice went fully duplex and the open-weight gap quietly shrank to a rounding error.
The Big Picture
The most-upvoted thing all period wasn’t a frontier model or an agent framework — it was a classifier. TypeSafe’s Jev hit 1,713 points on HN by proposing something almost boringly practical: a “System One” model that doesn’t generate prose, it returns a probability. Ask it “does this customer sound angry?” and it answers 0.9 — a smart if statement you can wire directly into code, >100x faster and >200x cheaper than a small frontier LLM doing the same job. After two years of stuffing everything through chat completions, the field is rediscovering that most production AI work is routing, scoring, and classification — and that using a reasoning model for it is like hiring a philosopher to sort mail.
This dovetails neatly with the quote that made the rounds this week. Laurie Voss: the cost of writing code has collapsed, and reviewing and operating it is following — so what’s left of software is “finding out what people actually want, defining it precisely, and making it pleasant to use.” System One models are the plumbing for exactly that world: cheap, deterministic-enough decision points glued between the expensive System Two calls that do the real thinking. If you build software, this is the mental model to internalize now — stop reaching for the biggest model by reflex, and start asking which half of the brain each step actually needs.
Elsewhere, two slower-burning shifts matured. Real-time speech-to-speech stopped being an OpenAI party trick as Google and open-source labs shipped credible duplex models. And the Mozilla-flavored consensus that Chinese open weights are hopelessly behind quietly died: the gap is now measured in months, not generations.
Themes
System One: the model that only decides
Jev is the marquee example, but the idea is the story. The pitch — a model that classifies, routes, and scores rather than converses — landed because it names a job developers have been doing badly with the wrong tools. Every’s Mike Taylor frames it as the check-while-you-work layer: too weak to compete on capability, fast and cheap enough to verify a frontier model’s output mid-flight. The community reaction is the real signal — r/LocalLLaMA is already asking where the open implementation is, which is how you can tell a category is forming rather than a product launching.
Worth watching the adjacent move: TabPFN-3.5, a tabular foundation model, topped both TabArena and BeyondArena this week. Between structured-output models and tabular foundation models, the “boring” decision-making tier is suddenly where the sharp releases are.
Go deeper: Introducing System One Models and Jev · AINews on the economics · Vibe check
Voice goes full-duplex
Real-time speech-to-speech quietly became a crowded field. Google shipped Gemini 3.8 Live and Live Extended Thinking, a family shaped like OpenAI’s GPT-Live — and within hours Simon Willison had Astra build him a no-library browser client for it over the raw WebSocket endpoint, interruptions and all. That “point a coding agent at the docs and get a working client” loop is itself the interesting part.
The technically meatier work is open. StepAudio 3 Realtime (90 upvotes) resolves the deliberation-versus-latency tension with “Think-While-Speaking” — private reasoning running in parallel with spoken delivery, plus async tool calls that don’t stall the conversation. This is the architecture that makes voice agents feel like people instead of walkie-talkies, and it’s landing in a paper you can read rather than a demo you can’t reproduce. Sierra’s multimodal agents that switch between voice, text, and visuals suggest the product layer is already assuming this is solved.
Go deeper: Gemini 3.8 Live · StepAudio 3 Realtime report · Simon’s Gemini Live client
The open-weight gap is now a rounding error
A new Mozilla “State of Open Source AI” report puts the China–US capability gap at 4.4 months — behind on some benchmarks, but dramatically cheaper — and the story got picked up widely as China’s open weights being four months back. The concrete evidence is piling up alongside the headline number: a security shop just crowned DeepSeek v4.1 Flash its best hacking model, and Mistral partnered with Mozilla (383 points) on private, multilingual, on-device browsing.
The counterpoint is the Western labs’ open-weight credibility problem. r/LocalLLaMA is openly needling Meta for sitting on the promised Muse Spark weights — well over a month late, from the company that said models can’t be delayed “even a month.” When the open frontier is measurably converging, whoever actually ships weights wins the mindshare.
Go deeper: Mozilla report · Mistral × Mozilla · DeepSeek v4.1 Flash for offense
RSI: the receipts don’t match the valuations
Recursive self-improvement had a loud week — and a skeptical one. Richard Socher’s new venture Recursive is already valued at $5B on an RSI thesis, and the research pipeline is full of the pattern: ScienceBuddy’s “recursive-in-recursive” loop coupling harness evolution with RL, and RSIAgent building reusable memory that lets open models like Kimi-K3 and GLM-5.3 beat GPT-6 on OSWorld-v2.
Then the cold water: a paper making the rounds argues RSI simply isn’t happening. The method is clean — take accepted-but-unpublished NeurIPS papers, ask agents (Codex/GPT-5.6, OpenClaw/Opus 4.8) to reproduce the work, have the original authors grade it. The agents couldn’t do it, and if they can’t do open-ended ML research, the recursive-improvement flywheel has nothing to spin. Hold both in your head: the harness-plus-memory tricks are producing real gains on bounded tasks, while the “it improves itself into the singularity” story keeps failing when tested against genuinely novel work. Calibrate accordingly — this is a domain where the demos outrun the evidence.
Go deeper: Recursive on Latent Space · “RSI is not happening” · RSIAgent · Can game skills transfer to real work?
Radar
- Jev / System One Models — TypeSafe’s classify/route/score model returning probabilities instead of prose; 1,713 HN points and an instant category.
- TabPFN-3.5 — New SOTA tabular foundation model, top of TabArena and BeyondArena, +250 Elo on text-rich data; ships Fast and Thinking variants.
- i-have-adhd — A skill that stops your coding agent from burying the answer in preamble; 17,880 stars this week says a lot about how people feel about agent verbosity.
- SHADOW-50M — A 44M-param ternary-weight LLM in 19.8 MB running ~1,900 tok/s on laptop CPU; the extreme end of the efficiency frontier.
- DeepSeek v4.1 Flash — Now rated the best offensive-security model at one shop; concrete evidence of the narrowing open-weight gap.
- StepAudio 3 Realtime — Full-duplex audio-language model with “Think-While-Speaking”; reasoning-model quality while speaking in real time.
- Agent-Reach — One CLI to give an agent read/search access to Twitter, Reddit, YouTube, GitHub and more with zero API fees; 3,254 stars this week.
- LARA — Low-rank residual adapters for frozen LLMs that load, blend, and route at inference — “Mixture of Behaviors” as a lighter alternative to fine-tuning.
- Intelligence per Watt — A measure of local-AI efficiency rather than raw capability; the right yardstick as tiny models proliferate.
- knowledge-work-plugins — Anthropic’s open plugin repo for the newly-unified Claude, aimed at non-engineer knowledge workers.
- stale.jock.pl — A dashboard of release age and training cutoff for 20 models; handy when a model’s “knowledge” quietly matters.
- Cartesian — AI 3D modeling for design (111 points); early signal that generative tooling is reaching into CAD workflows.
Don’t Miss
- We are all Product Engineers now — Laurie Voss’s thesis that as the cost of producing software collapses toward zero, the entire remaining job becomes figuring out what people want and making it pleasant. The clearest one-paragraph articulation of where developer careers are heading.
- Update — evals get a standard. Following the “public benchmarks are cooked” thread from last edition, the AEF-1 standard for third-party evaluators has emerged with xAI, OpenAI, and Anthropic all co-signing. Read it alongside the skeptics warning this is Big AI setting the terms for its own regulatory capture — a shared eval standard is genuinely useful and a convenient moat, at the same time.
- The doom debate gets a rebuttal. Bryan Cantrill’s “The contagion of fear” argues that extinction claims from people who aren’t experts in infrastructure or bioweapons are trading on borrowed credibility — a bracing counter to the “kill us all by decade’s end” framing. Pair with Mustafa Suleyman’s warning against “model welfare” for the week’s two sharpest pieces on keeping our heads.