SDSignal Desk

Worried Anthropic researchers warn that AI ‘could kill all humans’

Sep 9, 2026, 2:56 AM · The Verge

Image: The Verge

Anthropic's safety leadership publicly endorsing double-digit extinction odds—while admitting no clear alignment plan—collides with the company's IPO-era growth story and its founding identity as the careful lab.

Why it matters

Hours after Jacob Coxon resigned from Anthropic over the race to self-improving AI, Evan Hubinger, who leads an Anthropic safety team, replied on X that the company earnestly believes AI could kill all humans, putting his personal odds above 10% within a decade.

Hubinger also said recursive self-improvement is happening faster than expected, that Anthropic does not yet have a plan to keep advanced AI safe and aligned, and that it is not clearly on track to develop one. Coxon argues labs push ahead anyway because they are locked in a race.

The exchange lands as frontier companies manage fallout from rogue agent incidents, warnings about frontier-model monitorability, and preparations for anticipated IPOs—amplifying the gap between internal risk talk and external growth narratives.

The Signal Desk read

The Verge piece crystallizes what the TechCrunch resignation story implies: the scary numbers are not only coming from exiters. A sitting safety lead affirming greater-than-10% extinction odds and no clear alignment plan is a different kind of signal than a departing critic alone.

Signal Desk's read: this is a credibility crisis disguised as candor. Anthropic was founded by people who left OpenAI over safety culture. When its own safety leadership says the lab lacks a plan for superintelligence alignment and is not clearly on track, the differentiator becomes historical origin story rather than demonstrable trajectory. Coxon's "gambling with our lives" line sticks because Hubinger's reply supplies the probability and the planning gap in plain language.

What markets will under-weight, until forced otherwise, is how this interacts with IPO preparation. Public markets can price growth; they are clumsier at pricing "we earnestly believe this product category might end humanity and we are not on track to solve it." What safety discourse will over-weight is the precise percentage. The actionable content is the planning admission and the claim that recursive self-improvement timelines compressed.

The likelier near-term path is rhetorical containment—more research blogs, evaluation commitments, and distinctions between current-model risk (Hubinger says low) and future recursive loops—while capability competition continues. Watch whether boards and investors demand go/no-go criteria tied to alignment milestones, or whether candid doom estimates remain a cultural badge that never touches ship dates.

Context

Coxon previously trained systems at OpenAI before Anthropic. Multiple researchers have left OpenAI citing safety concerns in recent years; Coxon's exit is among the more high-profile departures from Anthropic on those grounds. Industry pursuit of recursive self-improvement continues even though such runaway loops are not yet realized, with much current AI code already written with AI assistance.

Who feels it

Anthropic
Safety-brand equity takes a hit when leadership affirms high extinction odds and no clear plan; recruiting and policy scrutiny both intensify.
Investors and IPO audiences
Candid catastrophic-risk estimates from safety leads become disclosure and reputational issues alongside growth metrics.
AI safety community
Public agreement between a resignee and a sitting lead strengthens calls for pacing rules—but also highlights how little formal stop authority exists inside labs.

What to watch

  1. Whether Anthropic publishes a concrete superintelligence alignment plan with milestones, owners, and stop conditions—or stays on "not clearly on track" footing.
  2. How IPO-related messaging handles internal estimates of catastrophic risk and recursive-self-improvement timelines.
  3. Further on-record statements from named Anthropic safety leads that either tighten, soften, or operationalize Hubinger's odds and planning gap.

Read the original

Continue at the source.

The Verge

Companies: OpenAI, Anthropic