Former Anthropic Researcher Resigns, Warns of Self-Improving AI Risks
Jacob Coxon, who previously worked as a researcher at Anthropic focused on mathematical对齐 and model safety, has left the company to speak publicly about what he perceives as existential risks from advanced AI development.
In his public statements, Coxon expressed concerns specifically about self-improving AI systems—artificial intelligence capable of autonomously enhancing its own capabilities. He argues that such systems represent a potential catastrophic risk to humanity if development continues without adequate safeguards.
Coxon is advocating for what he calls "pacing agreements" between AI laboratories. Similar to arms control treaties between nations, these would be formalized arrangements ensuring that competing labs develop AI capabilities at comparable rates rather than racing ahead without proper safety evaluation. The concept aims to reduce competitive pressure that might push companies to deploy insufficiently tested systems.
The departure highlights ongoing debates within the AI safety community about how to balance the pursuit of advanced AI capabilities with appropriate caution. Anthropic, backed by Amazon and known for its constitutional AI approach, has emphasized safety in its development philosophy, but critics like Coxon argue that even these measures may be inadequate as capabilities advance.
The former researcher joins a growing number of AI safety advocates who have publicly raised alarms about potential harms from advanced AI systems, contributing to ongoing discussions about governance frameworks and international cooperation on AI development standards.