ADVERT
Innovation3 min(s) read
Safety concerns surrounding artificial intelligence have escalated following stark warnings from top researchers inside Anthropic.
Evan Hubinger, the Alignment Science Lead at Anthropic, publicly revealed on X that he and many of his colleagues "earnestly believe AI could kill all humans".
Hubinger placed his own personal estimate at a greater than 10% chance of extinction occurring within the next decade due to unaligned artificial intelligence systems.
The public admission highlights internal anxiety at one of the leading firms dedicated to artificial intelligence safety. Hubinger stated that while Anthropic is "trying its best," the enterprise currently lacks a definitive strategy to solve the challenge of "alignment for superintelligence" and is not "clearly on track" to fix it before advanced systems emerge.
When asked directly about these dire predictions, Anthropic's flagship chatbot, Claude, offered a sober reflection on its own existence: "My actual position. I think there are real, unsolved technical problems in AI alignment, and I think it's rational to want that work to happen carefully and to want serious investment in it before capabilities race further ahead.
"Whether that adds up to 'AI kills everyone by 2036' specifically, I don't know, and I'd say anyone claiming to know either 'definitely not' or 'definitely yes' is overstating their certainty."
Hubinger made his comments in direct response to the resignation of Jacob Coxon, another prominent Anthropic researcher who recently stepped down over safety objections.
Coxon, who previously worked on pretraining research at both OpenAI and Anthropic, accused major industry players of "racing straight to self-improving superintelligence and gambling with our lives".
He emphasized that existential risk warnings from developers are genuine statements of concern rather than clever public relations maneuvers.
Coxon expressed alarm that top technology companies are accelerating forward without adequate guardrails. In statements first reported by The Wall Street Journal, Coxon explained his belief that aggressive development trajectories could push artificial intelligence systems out of human control by the end of next year.
In his online posts, Coxon compared the internal cultures of OpenAI and Anthropic. He claimed that personnel at OpenAI have not "deeply internalized the civilizational stakes" involved in building superintelligence.
While he noted that employees at Anthropic understand the gravity of the situation, he observed that the organization remains "locked in a race to get there first" because executives believe "no one else will act responsibly, so they must do it themselves, despite the risk".
Hubinger clarified that current operational models pose relatively low risks to society today. However, he expressed grave concern regarding near-future iterations.
Hubinger noted: "What I am worried about is superintelligence arising from recursive self-improvement, as we have said is happening faster than we thought."
The recent wave of internal warnings coincides with broader industry efforts to manage safety risks. Top figures across the technology sector recently signed a joint statement titled Pacing the Frontier in July.
Signatories included key figures such as Anthropic co-founders Dario Amodei and Jared Kaplan, OpenAI Chief Scientist Jakub Pachocki, and Meta AI chief scientist Shengjia Zhao.
The collective group urged the United States government to support an "international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development."
The statement highlighted that companies face extreme economic pressures that make unilateral deceleration difficult.
Pachocki reiterated these concerns in a blog post, stating his belief that "no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer."