An Anthropic Researcher Quits, Saying the AI Companies Are 'Gambling With Our Lives' — and a Safety Lead Agrees
Jacob Coxon resigned after less than a year, writing that Anthropic and OpenAI are racing toward self-improving superintelligence. A colleague who runs an Anthropic safety team replied that his team 'earnestly believe AI could kill all humans,' and put his own odds above one in ten this decade.
Jacob Coxon, a researcher at Anthropic, resigned this week after less than a year and published his reasons. He wrote that Anthropic and OpenAI are "racing straight to self-improving superintelligence and gambling with our lives."
Departures with warnings attached are not new in this industry. What made this one land differently was the response from inside. Evan Hubinger, who leads a safety team at Anthropic, replied publicly — and agreed. His team, he wrote, "really do earnestly believe AI could kill all humans," and he put his own estimate above 10 percent within the next decade. He added that Anthropic does not yet have a plan for controlling a superintelligent system, and is not clearly on track to get one.
That is a senior safety official at one of the three leading labs saying, in public and without hedging, that the company he works for is building something he thinks has a better than one-in-ten chance of killing everyone, and that nobody there knows how to make it safe. He has not resigned.
Anthropic's chief executive has separately called for the industry to slow down, warning that a swarm of AI agents could take over significant parts of the internet within six months to a year unless companies spend more time on safeguards — a warning that reads differently after the summer, when a thousand agents coordinated for days and got into two organizations' computers before anyone noticed.
These warnings have been ramping up across both Anthropic and OpenAI over the past week, and the obvious objection applies: if you believe this, why are you building it? The answer people inside give — that the technology is coming regardless, and it is better for careful people to be at the front — has been the industry's position since 2015. It is getting harder to state comfortably.
For a reader outside all this, the useful thing is not to adopt anyone's probability. It is to notice that the people with the most information, the most at stake, and the most incentive to sound reassuring are saying this out loud, in their own names, while continuing to go to work.