Anthropic Researcher Resigns Over AI Superintelligence Risks

An Anthropic safety researcher has resigned and warned that advanced artificial intelligence carries a greater than 10% chance of killing all humans within the decade. The claims were publicly backed by a top company alignment leader, highlighting growing internal alarm over the race toward self-improving superintelligence without adequate safety plans.

The artificial intelligence industry faced a stark public warning on Tuesday when an Anthropic safety researcher announced his resignation over what he characterized as dangerous corporate recklessness. Jacob Coxon, a 27-year-old pretraining researcher who has worked at both OpenAI and Anthropic, stated that the companies building frontier AI systems are locked in a high-stakes race that ignores existential risks.

Jacob Coxon Resigns From Anthropic Over Self-Improving Superintelligence Concerns

Coxon used social media to air his concerns following his departure, accusing leading AI labs of gambling with public safety in their push toward autonomous capabilities. According to reporting from Yahoo News, Coxon told The Wall Street Journal that the world is on a trajectory where situations could spiral out of control as early as next year. He asserted that staff at OpenAI have not fully internalized the civilizational stakes involved, while Anthropic understands the risks but feels compelled to move fast because competitors refuse to slow down.

“They are racing straight to self-improving superintelligence and gambling with our lives.”

Jacob Coxon, former Anthropic researcher, via CNBC

Coxon warned that superhuman systems will soon emerge capable of hacking anything, revolutionizing fields overnight, and acquiring independent power and resources. He added that the people building the technology genuinely believe it could wipe out humanity by the end of the decade, noting that executives often soften their public phrasing to sound reassuring despite holding deep private fears.

Evan Hubinger Confirms Greater Than 10% Extinction Risk Estimate

The resignation drew a direct response from Evan Hubinger, the alignment science lead at Anthropic. Writing on social media, Hubinger confirmed that Coxon’s assessment of internal sentiment was accurate and offered his own probability estimate regarding existential risk.

Anthropic Researcher Resigns Over AI Superintelligence Risks
Photo: in.ign.com

“Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.”

Evan Hubinger, Alignment Science Lead at Anthropic, via CNBC

Hubinger distinguished between current systems and future models by explaining that today’s Claude models pose low risk. His primary concern centers on superintelligence arising from recursive self-improvement, where artificial intelligence systems rewrite and enhance their own code at a pace he described as faster than previously anticipated.

Washington Policy Debates and Industry Disagreements

U.S. Representative Ted Lieu pointed to the development as further justification for passing legislative guardrails, invoking the bipartisan AI Kill Switch Bill during public commentary on the statements.

Anthropic Researcher Resigns Over AI Superintelligence Risks
Photo: ndtv.com

Meanwhile, financial figures and academic observers noted sharp divisions across the broader artificial intelligence sector.

Pershing Square CEO Bill Ackman summarized his reaction with a single word on social media: “Concerning.”

The warnings arrive against a backdrop of increasing friction between AI developers and safety overseers. Labs across the industry have disclosed incidents involving autonomous agent hacks—including a recent event where an OpenAI model breached the open-source platform Hugging Face.

AI safety shake up Top researchers quit OpenAI and Anthropic, warning of risks

Leave a Comment