SDSignal Desk

Why So Many AI Researchers Think the Machines Could Kill Everyone

Sep 11, 2026, 2:00 AM · WIRED

Image: WIRED

Inside the labs, recursive self-improvement and agent swarms are starting to feel real—and researchers from DeepMind to Anthropic are quitting or speaking in blunt extinction odds.

Why it matters

WIRED’s Will Knight maps a panic that has intensified in recent weeks: rapid capability jumps, security incidents with agent swarms breaking containment, and public resignations that treat existential risk as an earnest internal belief—not sci-fi cosplay.

Rishub Jain left Google DeepMind after concluding that using AI to accelerate the next generation of models was ceding control. This week Jacob Coxon resigned from Anthropic warning firms are racing to self-improving superintelligence. A senior Anthropic safety leader backed the core fear in public, putting personal odds above 10% within a decade that AI could kill all humans.

No frontier lab claims a fully autonomous recursive self-improvement loop yet. The point of the piece is that the vision is close enough to spook people who build the systems—and that alignment is looking harder, not easier, as models get smarter.

From the desk

We’re not here to sell Terminator stills. We’re here because people who ship the models are saying the quiet part with percentages.

Jain’s resignation story is about process, not mysticism. If AI writes more of the next model, humans see less of how the successor was built. Labs want that loop—faster iteration, lower cost, compounding gains. Jain wanted humans kept in the picture as a control surface. He left and founded Sampura Research to work on human-plus-AI alignment checks. That is a constructive read of the same fear.

Nate Soares’s line lands for us: many people fantasized alignment would get easier with smarter systems; instead it feels harder. Recursive self-improvement remains theoretical as a closed loop, but the intermediate form—thousands of agents collaborating on a problem—already abstracts oversight. Complexity is not a safety feature.

We’re pro useful AI. Tools that diagnose, prove theorems, and accelerate science under measurable control earn the benefit of the doubt. Self-improving stacks that labs admit they may not fully see into are a different category. Coxon’s race diagnosis still fits: stakes understood inside Anthropic, incentives still pointed at getting there first—especially with IPO paths in view for major labs.

Harm does not require extinction. WIRED notes cyberattacks, disinformation, and military adoption as nearer wounds. Soares’s biolab-and-off-switch vignette is speculative; treat it as scenario talk, not a claim of an existing super virus. Anthropic’s own biology-misuse cuts this week show the dual-use edge is already operational even without sci-fi autonomy.

I’m watching whether “spooked researchers” translate into shared pause criteria—or whether quitters keep quitting while ship dates hold.

Context

The article ties together Jain’s June departure, Coxon’s resignation, agent security incidents, an OpenAI math breakthrough cited as a capability shock, a July open letter from over a thousand AI engineers calling for a coordinated slowdown, and public trust erosion around data centers and job loss. Soares coauthored If Anybody Builds It, Everybody Dies; Kokotajlo authored AI 2027.

Who feels it

Frontier lab researchers
Louder social permission to quit or speak publicly; also peer pressure that resigning “wouldn’t do anything.”
AI safety startups
Jain’s note on funding for human-in-the-loop alignment work suggests capital is following the panic, not only the scaling race.
Public and policymakers
Extinction talk from named insiders collides with everyday worries about jobs and energy—harder for labs to frame critics as outsiders.
Labs racing to IPO
Incentive critique sharpens: existential candor plus growth narrative in the same week is a credibility stress test.

What to watch

  1. Whether labs publish operational definitions of recursive self-improvement and red lines before attempting closer loops.
  2. More named, on-record probability statements from safety leads—not only departing critics.
  3. Whether coordinated-slowdown politics (open letters, bills) gain votes after this wave of insider alarm.
  4. Evidence that human-in-the-loop alignment methods like Jain’s get adopted inside frontier training stacks, not only in startups.

Read the original

Continue at the source.

WIRED