Critical, Achieved
Three frontier labs shipped dedicated cyber models and safeguards in the same 48 hours — the exploit gap grew up into a product category. Meanwhile Fable 5.1 resets coding SOTA, the harness quietly becomes the thing that actually matters, and the citation layer underneath AI search turns out to be manufactured.
The Big Picture
The story we called “the exploit gap” last week stopped being a research anxiety and became a shipping category. On September 1, OpenAI announced that Astra is the first model to meet the “Critical” cybersecurity threshold under its Preparedness Framework — and shipped it anyway, wrapped in new safeguards. Within a day, Google DeepMind put out both Gemini 3.8 Flash Cyber and a proactive cyber-defense offering for governments and enterprises, while Anthropic leaned into enterprise frontier safeguards and alignment and security work. The disclosure-breaking capability we flagged is now sold as a defensive product, with the offense-side risk managed by policy rather than withheld. That’s the shift: the labs decided the way to contain a Critical-rated capability is to ship the blue-team version faster than the red team can weaponize it.
Underneath the model launches, a quieter and arguably more useful realization is landing: the model is no longer the variable that decides whether your agent works. A Show HN eval running nine harnesses against the same model found cost-per-pass varying 17x. The scaffold — permissions, memory, tool routing, context discipline — is now where the leverage lives. That reframes a lot of tooling decisions.
And Anthropic is, in the words of two separate reviewers, “so back.” Fable 5.1 is a genuine SOTA coding release, with a 75% cache price cut alongside 70% more output tokens — but the headline number is a jump on a brand-new science benchmark, not on the coding evals everyone already saturated. Which tells you where the frontier thinks the next hard problems are.
Themes
Cyber goes Critical — an update to the exploit gap
Last week the anxiety was that agents find bugs faster than disclosure can absorb them. This week the labs formalized it. OpenAI’s path-to-Astra post is notable less for the capability than for the governance choreography around it — a Critical rating met, then released with graduated safeguards, is the new template. Google is productizing both sides at once: an offensive-flavored Flash Cyber model and a defensive service.
The useful cold shower here comes from outside the labs: a security firm reports it found six real curl CVEs after OpenAI’s and Anthropic’s own agents came back with zero. Read the two stories together and the picture is honest — frontier cyber capability is real enough to rate “Critical” on a synthetic threshold and still miss what a tuned specialist harness catches. The bench number and the field result are diverging, which is exactly the pattern to watch.
Go deeper: OpenAI: Path to Astra · Gemini 3.8 Flash Cyber · DeepMind proactive cyber defense · Aisle vs. curl · Anthropic enterprise safeguards
The harness is the product
If you only change one mental model this week, make it this one. The FrontierHarness eval — same model, nine harnesses, 17x spread in cost-per-pass — is a small post with a large implication: your agent’s economics and reliability are now dominated by scaffolding, not weights. The corollary is showing up everywhere. AURA, an open-sourced Rust incident-response agent, makes the argument explicitly: permissions enforced deterministically outside the agent’s context so it can’t grant itself capabilities, workers scoped by domain, and — tellingly — “very good” root-cause accuracy on open-weights models. The intelligence was never the bottleneck; the containment and routing were.
This is also why persistent-memory and browser-control layers are trending as first-class infrastructure rather than features: agentmemory (519 stars this week) and chrome-devtools-mcp (738) are the plumbing that makes a mediocre model reliable. And on the maintainer side, Latent Space’s PRs NOT Welcome documents Vercel, Astro and tldraw replacing drive-by community PRs with “software factories” — teams of agents applying fixes — which is the harness idea applied to the open-source contribution pipeline itself.
Go deeper: FrontierHarness eval · AURA (Rust incident agent) · PRs NOT Welcome · LLMs: Intelligence vs. Cost · agentmemory
Fable 5.1 and the 180,000 lines nobody read
Fable/Mythos 5.1 is the real coding-model story of the period — reviewers report whole working projects from a single prompt at roughly half the tokens, and Simon Willison notes the suspiciously large jump on Terminal-Bench-Science (24.7% → 52.6%) against modest gains elsewhere. The frontier is moving toward long-horizon knowledge work, not more coding-bench points.
But the more consequential artifact this week is Rick Brewster’s account of shipping a from-scratch, clean-room reverse-engineered Direct2D rewrite for Paint.NET — 180,000 lines, written by Claude, “vibe coded,” essentially unreviewed. His own words: “I cannot possibly review 180,000 lines of code.” Set that against the two skeptical HN essays trending alongside it — AI Can Make You Suck Faster Too and The safest job from AI may be writing — and you have the central tension of the moment. The capability is real and the output ships; the review model that used to make code trustworthy has quietly been abandoned, and nobody has a replacement.
Go deeper: Every’s Fable 5.1 vibe check · Willison on Fable 5.1 · Paint.NET’s vibe-coded Direct2D · AI Can Make You Suck Faster Too
The citation layer is being farmed
Here’s an early signal worth watching closely, because it undermines a workflow many of us now depend on. An investigation found three sites generated 215,128 “best software” pages specifically to be cited by AI answer engines — and Perplexity cites them. Separately, an audit claims a third of Perplexity’s citations don’t actually contain the number they’re cited for. Put together: the retrieval layer that makes AI search feel trustworthy is simultaneously being gamed from the supply side and misattributed on the output side. If you’ve been treating cited answers as verified answers, this is your reminder that a citation is a link, not a fact.
Go deeper: Manufactured sources behind AI recommendations · Perplexity citation audit
Radar
- Gemini 3.8 Flash & Flash Cyber — Google’s fast tier plus a dedicated cyber variant; the Flash Cyber split is the notable move (335 HN points).
- Fable/Mythos 5.1 — new coding SOTA with a 75% cache price cut and 70% more output tokens; the pricing shift matters as much as the eval.
- DeepSeek-V4-Flash-Vision-Exp — vision support merged with Unsloth GGUFs already out; open-weights multimodal keeps pace.
- WebLLM + @huggingface/kernels — in-browser inference and 200+ WebGPU kernels landed together; local-in-the-tab is becoming real, and it pairs with reports of Apple caught off guard by AI demand for Mac Mini and Studio (495 points).
- AURA — open-source Rust incident-response harness with deterministic, out-of-context permission enforcement; a clean reference design for production agents.
- chrome-devtools-mcp — official Chrome DevTools for coding agents (738 stars/week); browser control as standard agent tooling.
- StudentSim — training LLM student simulators that mirror and update under tutoring; 369 upvotes, the period’s top paper, and a plausible eval substrate for AI tutors.
- SMELT — compute-matched scaling law for MoE looped transformers, saving 7–18% training FLOPs with gains largest on code (68 upvotes).
- agentmemory — benchmark-backed persistent memory for coding agents; the memory layer is consolidating into named projects.
- VoiceStudio — fully-local open-source ElevenLabs alternative (2,381 stars/week) covering cloning, dubbing and transcription in 646 languages.
- TimesFM — Google’s time-series foundation model trending again (681 stars/week), riding a broader real-time-forecasting push.
- Quasar 438B — a European frontier-scale entrant; worth a skeptical look given the marketing framing (126 points).
Don’t Miss
- How accurate have Ed Zitron’s AI-skeptic predictions been? — Dan Luu grades the most prominent AI skeptic against his own record. At 818 points and 944 comments it was the period’s biggest discussion by far, and it’s the rare backlash-adjacent piece that does the empirical work instead of vibes. Read it as calibration, whichever side you lean.
- Ethan Mollick — Agency and Agents — from the Hugging Face incident to “twilight factories,” a mental-model piece on where human agency sits once agents run the loop. Pairs well with the harness theme above.
- Willison on Claude’s published system prompts — Anthropic keeps publishing (and versioning) its consumer system prompts, and you can
.mdany doc page to diff them. A genuinely useful window into how a frontier lab actually steers behavior — including the newly aggressive anti-song-lyrics guidance.