Anthropic safety lead puts existential AI risk above ten percent this decade, admits no alignment plan exists
Evan Hubinger, who heads one of Anthropic's safety teams, publicly acknowledged that the company estimates the existential risk from artificial intelligence at more than ten percent over the next decade. In the same breath, he confirmed that Anthropic still has no concrete plan to ensure advanced systems remain safe and aligned with human values. The remarks came in direct response to the resignation of Jacob Coxon, a senior researcher who trained models at Anthropic and previously at OpenAI. Coxon posted on X accusing both companies of "racing straight toward self-improving superintelligence and gambling with our lives."
Coxon wrote that the people building the technology "truly believe it could kill us all by the end of the decade," yet the companies remain locked in a race that pushes them forward despite the risk. He described a "locked in a race" dynamic in which neither side is willing to slow down because the commercial advantage of reaching advanced systems first outweighs safety considerations. The departure is one of the most high-profile exits from Anthropic, a company founded by former OpenAI employees who left over precisely these safety concerns.
Hubinger did not deny the charges. He confirmed that the concern about self-improving superintelligence is "materializing faster than we thought," and agreed with Coxon's description that developers "truly believe" in the existential risk. His personal estimate: more than one in ten within a decade. The more troubling part of the response was the admission that Anthropic "still doesn't have a plan" for guaranteeing the safety and alignment of advanced systems, and that it is "not clearly on track" to develop one.
The dynamic Coxon described is not new to industry observers: models write code for the next generation of models, recursive self-improvement loops are a stated goal, and the pace is accelerating. But now the assessment comes from an internal safety-team lead, not external critics. The admission that there is no plan, and that the company is not on a path to produce one, turns a theoretical risk into a concrete operational gap.
The exchange arrives as both companies prepare for expected IPOs, against a backdrop of accumulating reports of "rogue agent incidents" and senior warnings about the ability to monitor frontier models. In recent years, additional researchers have left OpenAI for the same reasons. The cumulative picture is not of an isolated failure, but of an industry pushing capability faster than its capacity to control it — with safety personnel themselves acknowledging the gap is not closing.