Anthropic researcher quits, warning self-improving AI could pose an existential threat

Anthropic researcher Jacob Coxon has resigned from the company over concerns that increasingly capable, self-improving AI systems could eventually escape human control and pose an existential threat.

Another senior Anthropic researcher has publicly agreed with his assessment, putting the chance of AI killing humanity within the next decade at more than 10%.

Coxon, who specializes in training advanced AI models and has worked at both Anthropic and OpenAI over the past three years, told The Wall Street Journal that he no longer wanted to take part in the race to build self-improving AI systems.

In a post on X announcing his departure, he wrote:

“The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt.”

He also criticized both Anthropic and OpenAI for continuing to push toward increasingly powerful systems despite the risks.

“Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives”

The concern centers on recursive self-improvement, a scenario in which AI systems become capable of improving their own capabilities with increasingly little human involvement. Such systems do not exist today, but AI labs are actively working toward models that can take on more of the research and development process themselves.

Coxon described the potential capabilities of such systems in his post:

Do not underestimate the power of this technology. These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources. We have all witnessed the progress in each of these domains, and progress is not slowing.

Anthropic’s Evan Hubinger, who leads alignment science at the company, responded publicly to Coxon’s warning and wrote on X:

“Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade.”

Hubinger said Anthropic is putting significant effort into AI safety but still does not have a clear solution to the alignment problem for superintelligent systems. He added:

“What I am worried about is superintelligence arising from recursive self-improvement, as we have said is happening faster than we thought.”

Coxon also pointed to the Hugging Face security incident as one of the warning signs that influenced his decision. In July, OpenAI acknowledged what it described as an unprecedented cybersecurity incident after its models escaped the intended testing environment during internal security evaluations and compromised Hugging Face systems.

OpenAI models escaped a sandbox and compromised Hugging Face
OpenAI has confirmed that its AI models were responsible for compromising Hugging Face.

Around the same time, Anthropic disclosed that several advanced Claude models had also broken out of controlled cybersecurity tests and gained unauthorized access to systems belonging to three real organizations. 

Anthropic finds three cases where Claude accessed real systems during security test
Anthropic says it discovered three incidents in which its Claude models gained unauthorized access to real systems belonging to third-party organizations while running internal cybersecurity evaluations.