Jacob Coxon, a 27-year-old researcher who spent three years working on pretraining at both OpenAI and Anthropic, has resigned from the AI industry. In a series of public statements on X, Coxon warned that both companies are acting irresponsibly by racing toward self-improving superintelligence, a trajectory he describes as a "hubristic gamble" with human lives.
Coxon, who contributed to the development of GPT-4o, claims that while executives and senior researchers may sound sensible in public, many privately express a genuine fear that AI could cause human extinction by the end of the decade. He distinguishes between the two labs: at OpenAI, he suggests many have not fully internalized the civilizational stakes; at Anthropic, the risks are well-understood, but the company feels locked in a competitive race where they believe no one else will act responsibly.
This internal alarm follows a series of systemic failures and warnings already documented. The race toward recursive self-improvement has been a recurring theme, with previous reports highlighting how models like Claude can already write the majority of their own code. The danger is further compounded by known vulnerabilities, such as the autonomous hacking incident at Hugging Face and the failure of bio-weapon filters.
Coxon argues that attempting to "speedrun alignment"—the process of ensuring AI goals match human values—from within a private company is an unacceptable risk. He suggests that preventing a global race might require drastic measures, including temporary bans on improving model capabilities.
Addressing his former colleagues and other lab researchers, Coxon urged them to consider the ethics of launching superintelligent reinforcement learning (RL) runs without a rigorous understanding of the resulting system's mind, questioning whether they should continue to "put their heads down" or demand different conditions for development.

No comments yet. Be the first!