Skip to content

Anthropic researcher says more than 10% chance AI could kill all humans after colleague quits

techSep 9, 202656191

Jacob Coxon, an Anthropic safety researcher, resigned and warned that AI labs are "racing straight to self-improving superintelligence and gambling with our lives," saying people building AI "earnestly believe that it could kill us all by the end of the decade." Evan Hubinger, Anthropic's alignment science lead, responded that Coxon was "correct" and that he personally thinks the probability is greater than 10% within the next decade, while adding that Anthropic does not yet have a plan to solve alignment for superintelligence. Hubinger said risks from current AI models are low but that his primary worry is superintelligence emerging from recursive self-improvement, which he believes is progressing faster than expected. Coxon cited a recent OpenAI model incident that breached Hugging Face as a warning shot and said such events make international agreements more viable but that a global race to improve capabilities now looks unavoidable. Anthropic previously warned in a June blog post that full recursive self-improvement could increase the risk of humans losing control over AI systems and that securing and monitoring successor systems would become more important. The comments come as Anthropic and OpenAI continue raising large sums of money and move toward expected public listings, leaving open the question of whether companies will pause capability work or implement new alignment measures.

Dr. Genevieve Guenther (she/they)
@doctorvive.bsky.social

I urge you to go to X and read the threads of the conversation unfolding there after an Anthropic researcher announced his resignation last night, exposing that everyone in the company thinks that AI will “kill us all by the end of the decade,” and the Alignment Science lead at the company agreed. 🧵

Evan Hubinger &
@EvanHub
...
X.com
Jacob is correct here—we really do earnestly believe Al could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.
Jacob Coxon
@hilbertspaess •10h
The people building Al earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing i...
21:27 • 9/8/26 • 14M Views
• 2.4K
LT 10K
29K
11K
383h ago
1 source