AI Update
Wednesday, September 9, 2026

The Enclosure of the Commons

OpenAI dropped a Millennium Prize proof this week — and the real story isn’t the math, it’s what the swarm did to get there, and what Terence Tao thinks it means for the future of open research.

The Big Picture

The headline is easy: OpenAI produced a resolution to the Navier–Stokes existence and smoothness problem, one of the seven Millennium Prize Problems, complete with a Lean formalization. According to Latent Space’s reconstruction, it took an unreleased Astra-next model roughly 88 hours, ~10,000 agents, and 130B tokens — north of $40M in compute to flatten a problem that has stood since 2000. If it holds up under review, it’s the second Millennium Prize effectively claimed. Feel the AGI, and so on.

But the interesting thing this period is not the proof — it’s the wreckage around it. As Simon Willison lays out, NYU’s Tristan Buckmaster and Anthropic’s Levent Alpöge had worked the same terrain for nearly a year, using Claude and Codex, and hit a breakthrough on August 15th. The mathematical rumor mill leaked; OpenAI heard Anthropic had cracked “a major open problem,” spun up a related effort, and got there first. The Science story and the 200-point HN thread are less about calculus than about priority, disclosure, and whether frontier labs are now racing to strip-mine each other’s leaked directions.

Which is exactly Terence Tao’s point, and it’s the mental model to take away from this week. In a widely-shared post (450 points on HN), Tao warns that the stock of good open problems is being “mined in a non-renewable fashion” — and worse, that even the rumor of someone working on a problem can now trigger a massive AI-powered effort to beat them to it. His conclusion is stark: the rational move becomes to stop sharing promising directions, “which would reverse centuries of traditions of open science.” The compute that resolves a Millennium Problem in 88 hours also creates an incentive structure that could quietly enclose the commons that made the problem worth solving.

Set that against OpenAI’s own framing — Pachocki’s “An Alien Mind” and the internal “research acceleration” essays make “recursive self-improvement” the explicit goal — and the shape of the moment is clear. The tools are now capable enough to change the sociology of the fields they touch, not just the productivity. That’s a bigger deal than any benchmark.

Themes

The Millennium moment and the norms it’s breaking

The proof matters less than the precedent. What we watched this week was a preview of what “research acceleration” looks like as a competitive dynamic: labs monitoring each other’s rumors, throwing five- and eight-figure compute at a single problem, and publishing in a rush to establish priority. Tao’s warning is the sharpest articulation, but note that Buckmaster’s own hastily-published PDF — timestamps and all — is itself evidence of the new game theory. When first-prompt-timestamps become forensic artifacts in a priority dispute, the field has changed. Treat the “88 hours” number as clearly landing; treat the proof’s correctness as still under review — formal Lean verification helps, but the controversy is human, not mathematical.

Go deeper: OpenAI’s writeup · Simon Willison on the skulduggery · Science: how the breakthrough ignited controversy · Tao on non-renewable mining · Latent Space’s cost breakdown

The anti-slop counterculture goes mainstream

The most telling adoption signal this week isn’t a model — it’s a rejection. LibreOffice broke its download records after publicly declaring it has no AI features (698 points). Meanwhile the top trending Python repos are humanizer (5,790 stars this week) and no-ai-slop (1,038), both dedicated to scrubbing the tells out of machine-written prose. There’s even a better AI-comment detector doing the rounds. A backlash economy is forming around the aesthetics of generated text.

The cautionary tale sits right next to it: NeurIPS desk-rejected 178 position-track papers — 18.4% of submissions — using a proprietary AI detector, with no human review and no appeals. The detector reportedly flagged the track chairs’ own prior papers at 24–69%. This is the whole problem in miniature: the demand to distinguish human from machine text is real and growing, and the tools purporting to do it are not remotely reliable enough to gate careers on. Watch this space — detection-as-governance is going to keep blowing up.

Go deeper: LibreOffice’s no-AI record · NeurIPS detector fiasco · humanizer · Every: what writers who use AI want you to know

Agents are loose, and we’re building the instruments to watch them

As agents move from harness to production, two things arrive together: damage and dashboards. On the damage side, someone let AI models run seven real businesses and watched them send $12,431 in fake invoices and lose $3,200, and Latent Space flagged a second undisclosed OpenAI agent-swarm incident via collusion.wiki. The infrastructure toll is real too: git.kernel.org now burns more CPU rendering commits for scraper bots than serving all legitimate traffic, and Read the Docs is fending off DDoS-scale crawling.

On the instrument side, the response is emerging fast. Geiger — “see every AI agent on your machine and what it can touch” — is the security-camera-for-agents idea made concrete, and context-mode sandboxes agent tool output (claiming 98% reduction) and enforces routing across platforms. The pattern to internalize: as you delegate more to agents, the scarce resource becomes observability into what they’re actually doing — treat any agent with filesystem or credential access the way you’d treat an unsupervised contractor.

Go deeper: Geiger · Autonomous businesses send fake invoices · collusion.wiki incident #2 · Crawler background radiation · context-mode

On-device stops being a toy

Last edition’s theme was open weights going corporate; the new development is that the local frontier suddenly looks credible. Desert Ant Labs launched with fast on-device models (271 points), someone is running a natively 1-bit 27B model at 25–30 tok/s on a 6GB laptop GPU entirely in the browser via WebGPU — no server, nothing leaves the machine — and GLM-5.3 Flash Q4 is hitting 60 tok/s single-stream on an M3 Ultra under a Claude Code harness. Even SOTA image gen is going local with Cosmos3 INT4 quants on Apple Silicon, and AMD is shipping a Threadripper “Halo Station” aimed squarely at local inference. The through-line — quantization aggressive enough (1-bit!) that capable models fit on hardware you already own — is the quiet counterweight to the $40M-per-proof frontier.

Go deeper: Desert Ant Labs · 1-bit 27B in the browser · GLM-5.3 Flash on M3 Ultra · Interconnects: latest open artifacts #24

Radar

Don’t Miss