Anthropic Alignment Lead Puts Risk of AI Wiping Out Humanity Above 10%
Evan Hubinger says current models pose low catastrophic risk but fears self-improving superintelligence, as researcher Jacob Coxon quits Anthropic over the AI race.
Topics
News
- At Fintech Fest, PM Pushes For Global UPI Links to Cut Costs for Indians Abroad
- Paytm Plans AI Agent Push Beyond Payments
- Meta Launches Muse AI Agent to Handle Email, Shopping, Travel
- Anthropic Alignment Lead Puts Risk of AI Wiping Out Humanity Above 10%
- OpenAI Says AI Solved Decades-Old Maths Problem
- Temasek, Seraphim Back Pixxel in $100 Million Round
Anthropic Alignment Science Lead Evan Hubinger says he puts the chance of AI wiping out humanity within the next decade at more than 10%, after researcher Jacob Coxon quit the company over fears about the race toward self-improving superintelligence.
Coxon, who worked at OpenAI from 2023 until July before joining Anthropic, announced his resignation on Tuesday. He said neither company was acting responsibly.
“They are racing straight to self-improving superintelligence and gambling with our lives,” Coxon wrote on X.
Responding to Coxon, Hubinger said researchers at Anthropic “earnestly believe AI could kill all humans” and put his personal estimate of that happening within the next decade at more than 10%.
Hubinger said Anthropic was “trying its best” to address the risk but did not yet have a plan to align a future superintelligence and was not “clearly on track” to solve the problem.
He later stressed that he was not describing the danger posed by current AI systems.
“What I am worried about is superintelligence arising from recursive self-improvement, as we have said is happening faster than we thought,” Hubinger wrote.
Recursive self-improvement refers to AI systems taking an increasingly large role in developing more capable AI systems. Anthropic says AI is already accelerating parts of its development process, but that fully autonomous recursive self-improvement has not been achieved and is not inevitable.
Coxon said he was leaving after three years of pretraining research across OpenAI and Anthropic. He argued that competitive pressure was driving both companies toward increasingly capable systems despite concerns about their potential consequences.
“At OpenAI, many have not deeply internalized the civilizational stakes,” Coxon wrote. “At Anthropic, the stakes are well-understood, but they are locked in a race to get there first.”
The warnings come amid a broader debate inside the AI industry over whether safeguards can keep pace with rapidly improving systems.
A statement published in July and signed by 1,386 employees of frontier AI companies called for an international effort to develop technical and governance mechanisms that could deliberately slow frontier AI development if needed.
Anthropic’s Responsible Scaling Policy sets out safeguards intended to increase as model capabilities and associated risks grow. The company has also warned that recursive self-improvement could arrive sooner than many institutions are prepared for, while stressing that such a development is not inevitable.
OpenAI Chief Scientist Jakub Pachocki wrote on September 6 that no AI lab had solved alignment and monitoring well enough to continue scaling at maximum speed for much longer.
“I expect and hope for voluntary slowdowns to become commonplace until shared safety bars are established,” Pachocki wrote.


