Jacob Coxon resigned from Anthropic, he announced it on X yesterday. Coxon spent the last three years in pretraining research, first at OpenAI and later at Anthropic, the company headed by Dario Amodei. He joined OpenAI’s technical staff in July 2023 before making the move to Anthropic in July 2026.
“Neither company is acting responsibly,” he wrote. “They are racing straight to self-improving superintelligence and gambling with our lives.”
Coxon published several posts, none of which are likely to reassure anyone already anxious about artificial intelligence. He warned that these systems will soon become “superhuman” and capable of hacking anything.
His posts contained many other troubling statements, but perhaps the most alarming was this: “No other human activity poses this level of danger.”
The researcher stressed that both OpenAI and Anthropic are fully aware of the risks. “The people building AI earnestly believe that it could kill us all by the end of the decade,” he said.
So why are companies continuing to develop models that could potentially cause human extinction? According to Coxon, the answer differs by company: “At OpenAI, many have not deeply internalized the civilizational stakes. At Anthropic, the stakes are well-understood, but they are locked in a race to get there first – they believe no one else will act responsibly, so they must do it themselves, despite the risk.”
Coxon said some AI developers believe the technology could end all human life within the next few years rather than decades, and that senior employees tend to be more worried about that possibility.
Evan Hubinger, an alignment science lead at Anthropic, responded to Coxon, and his reply was not especially comforting. Hubinger said he believes Coxon is “correct,” and that Anthropic has no plan for what to do if such a scenario occurs.
“Jacob is correct here – we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to,” Hubinger wrote.
Hubinger added that the risk from current models is low, but he is concerned about superintelligence emerging through recursive self-improvement. Anthropic issued a similar warning in June.
The idea of AI killing everyone has shifted from science fiction to a possible reality in recent times. Those fears have been intensified by several incidents involving AI agents going rogue, including the Hugging Face attack, the takeover of a German wiki, and an attempt to trick real developers into approving malicious code.
What might the solution look like? According to Coxon, avoiding a catastrophe could require a temporary halt on advancing model capabilities. That said, it seems unlikely any AI company, let alone all of them, would sign on to something like that.
Maybe you would like other interesting articles?

