Jacob Coxon, a researcher specializing in training new AI models by consuming vast amounts of data, has resigned from Anthropic, expressing concerns that the company and its competitors are rushing to build AI systems they may not be able to control. Coxon, who previously worked at OpenAI, stated on X (formerly Twitter) that both companies are "racing straight to self-improving superintelligence and gambling with our lives." He joined Anthropic after working at OpenAI from 2023 until July 2026, where he contributed to GPT-4o research. His decision highlights mounting safety concerns within top AI companies regarding the industry's rapid advancements.
Coxon's most significant claim is that the fear of extreme AI risk is prevalent inside AI labs, not just an external criticism. He wrote, "The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt." He suggested that while executives and senior researchers might temper their public statements, they express greater fears privately. Evan Hubinger, Anthropic's Alignment Science Lead, also echoed these sentiments, stating he and his colleagues "earnestly believe AI could kill all humans" and that the company is not on track to align AI's goals with humanity's.
Despite the severity of his warning, Coxon did not advocate for a complete halt to AI development, expressing optimism for coordinated action. He pointed to events like the reported Hugging Face attack as potential catalysts for agreements between U.S. AI laboratories. However, he cautioned that voluntary cooperation might be insufficient to prevent a global race, suggesting that stronger measures, possibly including a temporary ban on improving model capabilities, might eventually be necessary. He urged lab researchers to consider the implications of unbridled AI development, questioning the wisdom of initiating a "superintelligent RL run without a rigorous understanding of its mind."