- Researcher Jacob Coxon announced he is leaving Anthropic after working on model pretraining at OpenAI and Anthropic over the past three years.
- Coxon accused both companies of irresponsibly racing to develop self-improving superintelligence and risking catastrophic consequences.
- Anthropic’s head of Alignment Science, Evan Hubinger, backed Coxon’s position and estimated a greater than 10% chance that AI could cause humanity’s extinction within the next decade.
Researcher Jacob Coxon announced that he is leaving Anthropic after spending the past three years working on model pretraining at OpenAI and Anthropic. He said neither company is acting responsibly enough as they race to build self-improving superintelligence, potentially putting humanity at risk.
“They are racing straight toward self-improving superintelligence and gambling with our lives,” Coxon said.
Coxon urged the public not to underestimate the capabilities of future systems. He said such systems could outperform humans in several areas, hack computer systems, dramatically accelerate scientific research and acquire significant influence in the real world.
He also said specialists working directly on frontier models are seriously considering scenarios in which continued development of the technology results in catastrophic consequences.
“The people building AI genuinely believe it could kill us all by the end of the decade,” Coxon said.
According to Coxon, the public statements of executives and leading specialists often sound considerably more cautious than their assessments in private conversations.
Coxon calls for voluntary constraints
Coxon drew a distinction between the approaches taken by OpenAI and Anthropic. He said many OpenAI employees have not yet fully grasped the potential consequences of superintelligence. Anthropic, by contrast, understands the scale of the risks but remains in the race because it fears that competitors will continue developing the technology regardless, he said.
Coxon described that approach as excessively risky. He argued that private companies should not have sole authority to decide whether to deploy systems capable of independently accelerating scientific research and their own development.
He said leading U.S. AI laboratories could potentially coordinate on safeguards. Possible measures could include voluntarily limiting the pace of development and, if necessary, temporarily refraining from further increases in model capabilities.
Coxon urged employees of AI companies to consider whether they are prepared to participate in training superintelligent systems without adequately understanding how those systems work or the risks they pose.
Anthropic alignment researcher responds
Anthropic’s head of Alignment Science, Evan Hubinger, responded to Coxon’s statement by backing his position. Hubinger put the probability that AI could cause humanity’s extinction within the next decade at more than 10%.
Hubinger also acknowledged that Anthropic does not have a ready-made solution for aligning superintelligence with human goals.
OpenAI chief scientist Jakub Pachocki previously highlighted the need to slow AI development in the absence of adequate safety guarantees. In an essay published on September 6, he said no laboratory had solved alignment and monitoring well enough to responsibly continue scaling systems at maximum speed for an extended period.
Source: Incrypted
