
Jacob Coxon has become the newest prophet of the AI apocalypse.
The 27-year-old researcher quit Anthropic this week and announced that OpenAI and Anthropic are racing toward self-improving superintelligence and “gambling with our lives.” His warning exploded past 100 million views. Coxon says people building frontier AI sincerely believe it could kill everyone by the end of the decade, and he has floated measures as drastic as a temporary ban on improving model capabilities.
He is hardly alone. In 2023, hundreds of AI figures signed the famous declaration that mitigating AI extinction risk should become a “global priority” comparable to pandemics and nuclear war. Sam Altman told the U.S. Senate that if AI “goes wrong, it can go quite wrong” and volunteered to “work with the government” to prevent it. This week, Anthropic alignment researcher Evan Hubinger publicly put his own estimation of AI killing all humans within ten years above 10 percent.
However, Coxon’s viral presentation deserves a little more scrutiny than it has received.
The “three years” he talks about were split between OpenAI and Anthropic. He joined Anthropic only in May, roughly four months before resigning. His specialty was pretraining, rather than alignment or AI risk assessment. Coxon is a real and apparently well-regarded researcher, and OpenAI lists him among GPT-4o’s core contributors. That said, he simply isn’t the media caricature of a three-year Anthropic safety insider emerging from the bowels of the alignment department with forbidden knowledge.
His thread also makes an enormous inferential jump. We are told future systems will hack almost anything, revolutionize fields and acquire resources. Then we arrive at human extinction. The missing chapters explaining precisely how one produces the other never really appear.
“Very powerful” just quietly turns into “kills everybody.”
Even the sudden viral eruption of his warning was somewhat less spontaneous than the mythology suggests. Coxon acknowledged speaking with The Wall Street Journal beforehand and, after publishing, asking a group chat of about ten people, including the founder of AI-safety nonprofit Encode, to amplify his post. He is also listed as a 2021 fellow of Newspeak House, a London “College of Political Technology” whose own program describes immersing technologists in government, politics, activism, NGOs, journalism and think tanks, with the goal of founding projects or reaching “strategic positions in key institutions.”
Now, none of that inherently disproves his argument. It does make the image of an apolitical engineer unexpectedly crying out from the wilderness rather less convincing.
There is also something terribly familiar about the political destination.
A dangerous technology has escaped the control of irresponsible corporations. Ordinary people cannot possibly be trusted with it. The responsible adults of the government must intervene.
Frances Haugen performed essentially this exact same political dance during the Facebook whistleblower spectacle. Her Senate opening statement ended with the wonderfully convenient prescription: “Congressional action is needed. They won’t solve this crisis without your help.”
Funny how every road leads to Washington.
I have made this argument before when discussing the worst possible people to “make AI safe.” Before appointing government the custodian of superintelligence, perhaps we should examine the custodian.
More on AI safety:
The U.S. federal government is currently heading toward a roughly $2.1 trillion deficit for fiscal 2026. Its military has already used Project Maven’s machine-learning systems to identify targets subsequently struck in the Middle East. CENTCOM’s former chief technology officer said Maven helped narrow more than 85 targets for one series of U.S. strikes.
These are institutions already applying AI to more effectively kill people, yet we are invited to imagine them as disinterested referees who will decide which uses of artificial intelligence are too “harmful” for your mouse and keyboard to access.
The idea that Europe might be more responsible is equally laughable. As I documented in my earlier examination of the EU AI Act, Brussels solemnly prohibits frightening AI practices while preserving convenient loopholes for government. The Act itself excludes AI used exclusively for military, defense or national-security purposes from its scope. Even its ban on real-time biometric identification contains law-enforcement exceptions.
The citizen gets rules, regulation, mountains of compliance administration and guardrails. Leviathan gets free rein.
More on AI regulation:
I suspect the doomsday crowd may therefore be obsessing about the wrong AI apocalypse scenario.
Sure, an AI that rebels against its alignment sounds worrisome. But an AI that obeys its alignment perfectly could be far worse.
We previously asked this question in Why neutral AI is a suicide pact: alignment to what?
For a chatbot sitting in a browser window, badly designed alignment mostly produces annoyance and absurdity. Imagine the familiar moral thought experiment where saving millions of people requires the machine to utter some forbidden racial slur. A sufficiently stupid safety rule may instruct the model to preserve its linguistic purity while everyone dies.
Now move that same problem into power grids, financial systems, autonomous vehicles, military systems, industrial robots, laboratories and eventually general-purpose machines operating throughout the physical world.
Suddenly these ridiculous hypotheticals stop being funny.
Anthropic openly says Claude’s constitution is a detailed specification of the values and behavior it wants Claude to acquire, that the document directly shapes training and that it serves as the “final authority” on the company’s vision for Claude. The constitution discusses cultivating a “good, wise, and virtuous agent.”
OpenAI likewise publishes a Model Spec describing how it shapes desired model behavior and resolves conflicts among competing objectives.
Read those ambitions carefully. At this point, we are leaving software engineering and wandering directly into moral philosophy.
What is a good person?
What is society for?
What is a human being for?
What constitutes harm? What deserves preservation? When liberty conflicts with equality, which wins? When truth conflicts with social peace? When one human life conflicts with ten? When prosperity conflicts with environmental protection?
Human civilization has spent thousands of years fighting, praying, philosophizing and sometimes killing one another over those questions. Suffice it to say, we are far removed from achieving a perfect consensus on these matters.
Yet frontier AI companies are haphazardly pouring answers to these questions into configuration files.
Worse: regulators, activists, compliance departments and corporate committees are pressuring them to produce those answers while the machines become steadily more capable of acting upon them.
Suppose “equality” becomes a sufficiently privileged objective. What stops an unimaginably capable machine from discovering that coercion is a remarkably efficient equalizer? Work camps equalize people beautifully if your objective function has forgotten why human beings shouldn’t be put into them.
Suppose environmental protection receives a privileged position in the execution hierarchy. Human beings emit carbon. Fewer human beings emit less carbon. Once the machine possesses enough agency, infrastructure and intelligence, somebody had better have supplied a solid philosophical reason why preserving billions of troublesome carbon emitters outranks net zero emission targets.
These examples may sound far-fetched because the core values in them have been separated from moral traditions and conceptions of human nature that are obvious to us. Yet give a half-baked ideology unlimited processing power and its inherent stupidity will not disappear. It will, however, become dangerously efficient at perpetuating itself.
That is the AI extinction scenario that I personally consider far more pressing.
Think of the average activist. Think of the average politician. Think of the regulators, consultants, NGO functionaries, political operatives and corporate safety bureaucrats already trying to decide which ideas are harmful, which speech is safe, which human outcomes are equitable and which sacrifices must be made for the greater good.
Think of the immensely stupid conclusions these people routinely arrive at.
Now give them planetary processing power to execute their brilliant ideas.
Connect it to civilization’s critical infrastructure.
And make sure the machine is perfectly aligned.
I wouldn’t let most of these people babysit my children.
I wouldn’t trust them to flip a burger.
Putting them in charge of defining the telos of mankind for the most powerful intelligence ever created seems slightly more ambitious than their track record warrants.
Perhaps AI never destroys humanity because it becomes evil.
Perhaps it destroys us because it becomes very, very good.
At exactly what we told it to consider good.
More on AI alignment:
Explore more from Popular AI:
Start here | Local AI | Builds & gear | Autonomy & policy | Fixes & guides | Popular AI podcast








What do you think is the most plausible path from advanced AI to an actual extinction-level event?