Anthropic researcher quits the AI industry, citing 10%+ extinction risk — colleague backs him up
Key Points
- Anthropic researcher Jacob Coxen quit the AI industry citing extinction risk above 10 percent, drawing support from alignment scientist Evan Hubinger, who shares the same probability estimate.
- Hubinger stated Anthropic and peers lack a plan to solve alignment for superintelligence and are not on track to develop one, undermining confidence in safety measures.
- Critics challenge the extinction claims as lacking mechanistic rigor, while a separate concern emerges that coordination between AI labs on safety and pacing could trigger antitrust violations.
Summary
An Anthropic researcher quit the industry this week citing extinction risk estimates above 10 percent, drawing both support and sharp criticism from peers and observers.
Jacob Coxen, who spent three years doing pre-training research at OpenAI and Anthropic, announced his departure in a post that received roughly 100 million views and 600,000 likes. His complaint was direct: "Neither company is acting responsibly. They are racing straight to self improving superintelligence and gambling with our lives."
Evan Hubinger, an alignment scientist at Anthropic with prior roles at MIRI, OpenAI, and Google, backed Coxen's concerns. Hubinger stated he personally believes the probability of AI-driven human extinction within the next decade is "above ten percent" and that "we do not yet have a plan to solve alignment for superintelligence and are clearly not on track to."
The post drew mixed reactions. Martin Casado, an Andreessen Horowitz partner who worked on thermonuclear weapons projects, compared the dissonance unfavorably—suggesting actual weapons work felt less bizarre than current AI development. On the other side, Nick Carter argued that if Coxen genuinely believed extinction was imminent, public resignation would be inadequate; he'd need to commit "acts of violence, sabotage, and terrorism" to match his stated beliefs. Eliezer Yudkowsky offered qualified support, saying "thank you for saying plainly what you believe."
The credibility question
Critics pointed out that a 10 percent extinction probability claim lacks the detailed mechanistic support needed to carry weight. Derek Davison called for clarity: "If you publicly assign a greater than 10% probability to human extinction within a decade, you owe people a clear explanation of how you reach that number plus what would change your mind." Hubinger did link to Anthropic's published risk report, which discusses scaling policy and pacing frontier development with other labs, but observers questioned whether the extinction framing amounts to "vibe prediction" rather than rigorous analysis.
Regulatory collusion risk
A separate tension emerged around the labs' stated desire for external regulation. If Anthropic, OpenAI, and other AI leaders coordinate on safety standards and pacing agreements, they could face antitrust exposure—coordination that looks like collusion from a legal standpoint. The hosts noted this could become textbook antitrust violation, even if the labs' safety intentions are genuine. The concern is that regulatory capture, while a potential side effect, is being used by observers to dismiss calls for regulation wholesale—when in fact some outside voices are now accepting that risk as worth the safety benefits.
Every deal, every interview. 5 minutes.
TBPN Digest delivers summaries of the latest fundraises, interviews and tech news from TBPN, every weekday.