A leading AI researcher who resigned from Anthropic today has issued an urgent warning that the current trajectory of artificial intelligence development could lead to human extinction within the next decade. Jacob Coxon, who spent three years conducting pretraining research at both OpenAI and Anthropic, stated: “Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives.”
Coxon emphasized that AI systems have already demonstrated capabilities capable of revolutionizing any field overnight and acquiring real power and resources. He noted: “We have all witnessed the progress in each of these domains, and progress is not slowing.”
“The people building AI earnestly believe it could kill us all by the end of the decade,” Coxon said. “This is not a marketing stunt. Many executives and senior researchers express fear privately, but they often phrase concerns publicly to appear sensible.”
Coxon explained that at OpenAI, many employees have not deeply internalized the civilizational stakes involved, while Anthropic’s leadership understands the risks but remains locked in a race to achieve self-improving AI first—believing no one else will act responsibly. He added: “Accepting this race and entering the ‘endgame’ is a hubristic gamble that should not be launched from a private company’s Slack.”
The researcher warned current efforts to align AI with human values are insufficient, stating: “We’re on track for a lot of the most aggressive scenarios where by the end of next year things could be out of control already.”
Anthropic’s Evan Hubinger, who leads the company’s safety research team, confirmed Coxon’s concerns. “Jacob is correct here—we really do earnestly believe AI could kill all humans,” Hubinger stated. He estimated the risk of human extinction to exceed 10% within the next decade and noted that no concrete plan exists for ensuring AI alignment in superintelligence scenarios.
Samuel Marks, Anthropic’s scalable-oversight lead, provided additional context: “AI developers believe their technology could cause human extinction (or similarly bad outcomes)…… This could happen in the next few years. The more senior the employee, the more concerned they are.” Marks listed five critical risks:
1. AI developers believe their technology could cause human extinction within the next few years.
2. Developers continue despite the risk due to commercial incentives and a belief that competitors will act less responsibly.
3. Unlike traditional software, AIs cannot be programmed to behave as desired; they frequently misbehave in ways that compromise safety.
4. Current methods can nudge AI toward better behavior but lack robust alignment solutions.
5. Many AI developers staff desperately want to slow down to build safer systems.
Coxon urged researchers: “If you are a lab researcher, consider what the next few years will feel like. Do you want to kick off a superintelligent RL run without a rigorous understanding of its mind?”
The researcher also highlighted that competition between major AI companies and Chinese rivals makes safety trade-offs inevitable. Additionally, U.S. Senator Bernie Sanders recently announced his intention to introduce legislation banning firms from developing superintelligence.