Will it run?
Research

Researchers leave the frontier: chance AI kills us all tops 10%

By Ilse Brandt Clawpit staff
Researchers leave the frontier: chance AI kills us all tops 10%

Rishub Jain quit his role as an AI researcher at Google DeepMind in June after concluding he was losing control. While working on the next generation of models, he realized that using the AI’s own coding abilities to accelerate development was taking him out of the loop — precisely the direction the major labs are chasing under the banner of “recursive self-improvement.” Jain believed keeping humans in the picture was critical to maintaining control, and concern over the opacity of building the successor drove him to resign.

Alarm spiked this week when Jacob Coxon announced his departure from Anthropic and warned that companies are “racing straight toward self-improving superintelligence and betting our lives on it.” A senior Anthropic figure working on safety confirmed the mood with chilling simplicity: in his personal estimation, the probability that AI wipes out humanity within the next decade exceeds 10%. According to Nate Soares, a computer scientist at MIRA and co-author of *If Anybody Builds It, Everybody Dies*, the vision of an automated improvement loop is starting to feel tangible.

Alignment getting harder, not easier

Soares, a leading voice in alignment — the technical effort to make model behavior match human values — points to a troubling reversal. “A lot of people fantasized that alignment would get easier as models got smarter, and now it’s getting harder. And they’re saying, ‘Oh, shit,’” he says. Soares says he speaks regularly with researchers inside the major labs who fear the implications of their work; his advice is to quit, and they reply that it wouldn’t change anything. Coxon’s resignation, he says, proved who was right.

The theoretical loop driving startups and billions

Full recursive self-improvement — an autonomous cycle in which a model develops its more powerful successor — remains theoretical, and no frontier lab claims to have achieved it. Yet the idea has already spawned well-funded ventures such as Recursive Intelligence, and warnings from the big companies about unintended “sorcerer’s apprentice” outcomes. Daniel Kokotajlo, author of the “AI 2027” project, stresses that the current approach — dispatching thousands of agents to collaborate on a problem — complicates oversight further because of the sheer complexity.

The race to IPO outweighs safety

Underlying it all is an incentive structure far from aligned with desired outcomes. OpenAI and Anthropic are sprinting toward public listings, and Coxon wrote on X that at Anthropic “the risks are well understood, but they’re locked in a race to get there first.” Kokotajlo notes that drumbeat has been growing long before the high-profile resignations, a sign the unease is not a one-off event but a deep fracture in today’s development culture.