We've heard these kinds of warnings before but this week there was a new one which really got a lot of attention, in part because the person making it seems to have a credible position from which to judge. His name is Jacob Coxon and on Tuesday he resigned his position at Anthropic out of concern that he was helping build something very dangerous.
Do not underestimate the power of this technology. These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources. We have all witnessed the progress in each of these domains, and progress is not slowing.
— Jacob Coxon (@hilbertspaess) September 9, 2026
A common response is “if they truly believe this, why are they still building it?” At OpenAI, many have not deeply internalized the civilizational stakes. At Anthropic, the stakes are well-understood, but they are locked in a race to get there first - they believe no one else…
— Jacob Coxon (@hilbertspaess) September 9, 2026
I am optimistic about the potential for coordination. Warning shots like the Hugging Face attack have made pacing agreements between U.S. labs more viable. I don’t feel like we’re on track to prevent a global race, which may require costly actions such as a temporary ban on…
— Jacob Coxon (@hilbertspaess) September 9, 2026
His use of the word "alignment" refers to the effort required to make sure AI systems have goals that are aligned with human wishes, i.e. not goals they make up for themselves and not goals that would be harmful to people.
He wrapped up his thread with a plea to other AI insiders to consider taking a different path now.
If you are a lab researcher, I urge you to consider what the next few years will actually feel like. Do you want to kick off a superintelligent RL run without a rigorous understanding of its mind? Should you put your head down because “it’s happening anyway” - or take this moment…
— Jacob Coxon (@hilbertspaess) September 9, 2026
Coxon got some support from another AI researcher at Anthropic.
Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to. https://t.co/QAIHiFP3QZ
— Evan Hubinger (@EvanHub) September 9, 2026
He clarified that he doesn't think anything out now could kill us all. The real concern is the speed at which current models could be used to improve future models.
Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to. https://t.co/QAIHiFP3QZ
— Evan Hubinger (@EvanHub) September 9, 2026
Using AI to write better AI systems is already happening and it means the rate of improvement could actually accelerate dramatically. You arguably already have models that are as capable as grad students at a range of activities, including math and programming. What happens when AI can exceed even the leading experts? In fact, there was some fresh evidence this week on that front.
OpenAI said on Tuesday that its newest artificial intelligence technology had solved one of the “Millennium Problems,” a collection of important unanswered math questions meant to push the world’s leading mathematicians to new heights...
Over the past year, A.I. systems successfully solved a wide range of problems that have bedeviled mathematicians for decades. But these problems were not as complex, nor as closely watched, as the one that OpenAI’s technology has solved over the past several days. The Millennium Problems are among the most heavily researched in the field.
“This is a spectacular culmination of the arc we have seen over the past twelve months,” OpenAI researcher Sébastien Bubeck said of the company’s new solution.
The company announced that one of its latest models, which has not yet been released to the public, needed just 88 hours to solve what mathematicians call “the Navier–Stokes existence and smoothness problem.” This problem involves a series of equations that are often used to predict the weather...
OpenAI said that it solved the Navier-Stokes problem using as many as 10,000 “A.I. agents” working in concert. Running such a large number of A.I. systems is likely to have cost millions of dollars, a result of the enormous amounts of electrical power needed to operate the specialized chips that drive A.I. technologies.
Getting back to the warning, OpenAI used a model that is not yet publicly available to do this. It's the cutting edge right now. But six months from now, this could be available to everyone and something much smarter could be solving new problems in a lab somewhere.
Coxon's story went so viral, thanks in part to it being picked up by Bernie Sanders and others, that there is now some question about whether all of this was being coordinated.
This whole Jacob Coxon story is really wild… 🤔🤷♂️ pic.twitter.com/TGwSUhReqw
— John Ziegler (@Zigmanfreud) September 10, 2026
Elon Musk said this:
Seems like a setup
— Elon Musk (@elonmusk) September 10, 2026
There could be some truth to it but keep in mind that Musk himself has warned many times that there is a bad path AI could take if we're not careful.
As for Sanders, he has been pushing for more AI regulation for at least a year now. So the fact that he jumped on this doesn't necessarily mean he was behind it or knew about it in advance. He may have just seen an opportunity to amplify something that fit into his preexisting views about the dangers of AI.
That said, there's no doubt that those views do seem to have a partisan aspect to them these days. The issue of data centers has even played a role in some midterm races with people on both sides of the aisle seemingly against building them and Democrats trying to capitalize on their unpopularity.
Ultimately, I think what matters is whether Coxon has a point that these systems are dangerous or will be in the not too distant future. It's not just the machines themselves that are a threat, it's what our enemies might be able to do with them.
In a lengthy new report out Thursday, Anthropic said that criminal hacking gangs, Chinese security bureaus, Russia-linked spies and Yemen-based arms manufacturers, among others, have exploited a web of fraudulent accounts and online services to access some of the company’s advanced AI assistants, known as Claude.
And they’re using those tools to attempt attacks or design weapons that previously lay beyond their capabilities.
In the last eight months alone, Anthropic said, people in countries where Claude should be restricted circumvented those controls to automate cyberattacks, try designing advanced military hardware, spread propaganda at scale and mount broad digital surveillance campaigns. Though Claude models are broadly available commercially, Anthropic tries to restrict access in major U.S. adversaries, including Russia, China and Iran.
The AI giant also said it uncovered five cases this year where scientists in unspecified foreign countries used their models to research dangerous pathogens, work that may have been conducted as part of a bioweapons program.
So forget about the science fiction part of this, i.e. rise of the machines acting on their own. The threat of China or Russia using our own advances to hurt us are real enough and are already happening now according to this report. Maybe we need to put more effort into stopping them before the tools themselves are powerful enough to amplify their bad intentions into real world damage.
