Read Time: 13 minutes
TL;DR
On September 9, 2026, a pretraining researcher named Jacob Coxon resigned from Anthropic — after three years across OpenAI and Anthropic — and posted a thread warning that the labs are “racing straight to self-improving superintelligence and gambling with our lives.” It has been seen tens of millions of times. Buried in it is a line that is not abstract to me at all: these will soon be “superhuman systems that can hack anything.” That is my beat. I am not an alignment researcher and I do not trade in p(doom), but I have spent 2026 documenting the concrete, boring, already-here version of exactly what he is gesturing at — the model as attacker, agents that act with your credentials, autonomous offense. The safety people and the security people are describing the same animal from opposite ends. This is my attempt to translate his warning into the language of the security chair — taking it seriously without catastrophizing, presenting the case against it fairly, and landing where I always land: the facts don’t need the doom framing to matter, and fear is not a security control.
A note on what this is. This is an opinion piece, not a threat model. I am a security practitioner, not an AI-safety researcher, and existential risk is a genuinely contested debate among serious people. I will give Coxon’s argument its due, give the skeptics theirs, and be clear about which parts are my own read. Nobody in this piece is a villain, and I am not telling you what to believe about the end of the world — only what this looks like from where I sit.
What happened
Jacob Coxon spent three years doing pretraining research — the deep end of the pool — first at OpenAI, then at Anthropic. On September 9 he quit, publicly, and wrote a thread that has since been viewed tens of millions of times. The core of it is blunt: “Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives.”
He goes further. “The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt.” He describes executives who soften their phrasing for the press while privately expressing fear, and researchers at his own former employer who understand the stakes but feel “locked in a race to get there first.” His proposed remedy is uncomfortable and specific: public dissent from lab researchers, and — as a serious option — “a temporary ban on improving model capabilities” to buy time for international coordination.
What made me sit up was not the existential framing, which I have heard before. It was that this was not an outsider or a pundit. And it did not stay uncorroborated: Evan Hubinger, Anthropic’s own Alignment Science Lead, publicly agreed that Coxon was right about the fear inside the labs, and volunteered his own number — “I personally think it is >10% within the next decade” — while making clear he considers the risk from present-day models low, with his concern aimed at a future superintelligence. Hold onto that last distinction. It is the whole argument in miniature, and I will come back to it.
And Hubinger was not the only one. Samuel Marks, a scalable-oversight lead at the same company, added that “this could happen in the next few years.” A caveat that matters for fairness: as I write this, these are individual researchers speaking for themselves on social media, not an official Anthropic statement. But when this many people at one lab say the quiet part on the record within hours of each other, the pattern is itself the news.
The one line that is my job
Strip the thread down and one phrase is doing the heavy lifting for a security audience: these will be “superhuman systems that can hack anything… and acquire real power and resources.”
To most readers that is science fiction. To me it is a Tuesday with the dial turned up. I have spent this year writing about the un-fictional, already-shipping leading indicators of precisely that capability:
- Models being turned into the attacker rather than the target — which is the entire subject of When the Model Is the Attacker, my read of the Hugging Face incident.
- Agent skills and connectors weaponized into a supply chain that executes with real privilege.
- Autonomous agents pointed at the open ocean of public data to track and predict people, unattended, which I walked through in the maritime OSINT piece.
- Open-weight models you can backdoor or simply not trust, where the provenance of the thing running your infrastructure is itself the risk.
None of these is superintelligence. Every one of them is a rung on the ladder Coxon is pointing at from the top. “Can hack anything” is not a phase change that arrives one morning; it is the far end of a curve I have been plotting point by point all year.
He cited my beat, by name
Here is the detail that pulled this from “interesting” to “I have to write about this.” In the same thread, Coxon names a specific event as a reason for cautious optimism about coordination: “Warning shots like the Hugging Face attack have made pacing agreements between U.S. labs more viable.”
That warning shot is the one I dissected in July. From inside the AI-safety frame, the Hugging Face incident is an abstract data point in an argument about lab pacing agreements. From inside the security frame, it was a concrete thing that happened, with a mechanism, a blast radius, and a set of controls that would have changed the outcome. Same event, two vocabularies. The safety community reaches for it to argue about governance; the security community reaches for it to argue about logging, provenance, and least privilege. We are both right, and we are mostly not in the room together.
That gap — between the people modeling the mind of a future superintelligence and the people modeling the attack surface of the system shipping this quarter — is the actual subject of this post. Because the second group has something the first group needs: evidence.
Two tribes, one elephant
The AI-risk conversation has largely split into two tribes that talk past each other.
The safety tribe argues in the abstract and the future tense: alignment, mesa-optimizers, recursive self-improvement, the mind of a system that does not yet exist. Their strongest move is that they are reasoning about the thing before it can hurt you, which is the only time reasoning helps. Their weakest move is that abstraction is unfalsifiable and easy to dismiss, and it asks you to be frightened of a capability you cannot yet see.
The security tribe — my tribe — argues in the concrete and the present tense: this exploit, this connector, this leaked key, this incident last week. Our strongest move is evidence: we can show you the thing, reproduce it, and measure it. Our weakest move is that we are so busy with this quarter’s fire that we rarely lift our heads to ask where the curve goes.
Put the two together and you get something neither has alone. The safety people supply the trajectory; the security people supply the receipts. Coxon’s “systems that can hack anything” is a claim about the trajectory. My year of write-ups is a stack of receipts showing the trajectory is real and pointed the way he says. You do not have to accept his timeline or his p(doom) to notice that the leading indicators are not zero, and that they are getting stronger, not weaker.
The honest case against him
I promised the skeptics their due, and they have a real case — I would be a bad practitioner if I only steel-manned the alarm.
First, dramatic capability claims are not disinterested. “Our technology is so powerful it might end the world” is, awkwardly, also the greatest sales pitch and the strongest regulatory moat ever written. A lab that convinces governments only it can be trusted to handle world-ending power has argued itself into an incumbency no competitor can dislodge. Doom talk and market power point the same direction, and that should make you read every capability claim — including the scary ones — with one eyebrow raised. It is not a fringe objection: plenty of the reporters covering this very resignation noted that critics accuse the labs of hyping their products. Hyping the danger is a way of hyping the product.
Second, the honest members of the safety camp keep saying the quiet part: today’s models are low risk. Hubinger said it in the same breath as his ten-percent figure. The distance between “a chatbot that still miscounts letters” and “a system that can hack anything and seize resources” is not a rounding error; it is the entire disputed question, and confident extrapolation across it is exactly the move skeptics are right to challenge.
Third, there is an opportunity cost to apocalypse. Every hour the discourse spends on a hypothetical 2030 superintelligence is an hour it does not spend on the mundane harms that are here now and provably hurting people — fraud, non-consensual imagery, model-enabled crime, the concrete things I write about. A reasonable person can believe the far tail is overblown and that the near-term security reality is under-served. I am close to that person.
Where I actually land
So do I think a superintelligence kills us all by 2030? I don’t know, and — this is the point — I don’t need to, to do my job.
Here is the move I want to offer, because it is the one thing a security practitioner can contribute that a philosopher cannot: you can decouple the action from the eschatology. Whether p(doom) is one percent or thirty, the rational security posture in front of you is identical. Instrument the agent layer so you can see what these systems do. Treat model provenance as a supply-chain problem, because it is one. Assume the text your systems read is hostile, because it demonstrably is. Keep a human in the loop on irreversible actions. Design for least privilege as if the model will be turned against you, because this year it sometimes was. Every one of those is worth doing if Coxon is completely wrong. Every one of those is worth doing if he is completely right. That is what a good control looks like — it pays off across the whole range of the disagreement.
Coxon’s real contribution, for my community, is not the timeline. It is the reminder that the curve has a top, and that the people closest to the summit are frightened enough to walk away from a great deal of money and status to say so. You can discount their forecast and still take their fear as data. Insiders defecting is itself a signal, the same way I treat any credible insider warning about any system: not proof, but a reason to look harder.
And fear, on its own, is not a security control. It never has been. The useful response to a frightening trajectory is not to panic and not to look away — it is to build the instrumentation, the provenance checks, and the least-privilege boundaries that hold up whether the scary version arrives in three years or never. That is unglamorous, it does not trend on X, and it is the only part of this whole debate I can actually hand you on a Monday morning.
And there is a version of this future worth being optimistic about, which I do not want to lose in the alarm. The “superhuman” in Coxon’s sentence is a system that outgrows us. The superhuman I want is a person — an analyst, a defender, a builder — amplified by tools they own and understand, which is the case I made in AI Must Make Superhumans, Not Unemployed. You can hold that optimism and still log your agents; done right, the two are the same project.
Take the warning seriously. Take the skeptics seriously. Then go log your agents.
Stay paranoid. Ground the fear in evidence. Build the controls that pay off either way.
- X (Twitter): @SimonRoses
Further Reading:
- Jacob Coxon’s resignation thread (X, September 9, 2026)
- When the Model Is the Attacker
- How to Weaponize AI Agent Skills
- Tracking the Fleet of the Rich: Maritime OSINT, a Radio, and an AI Agent
- Do Open Weight Models Dream of Tokens?
- The Death of the Job: How AI and Robots Will Rewrite Work in the Next 10 Years
- AI Must Make Superhumans, Not Unemployed
Questions or feedback? Reach out via:
- Website: vulnex.com
- AI Security Strategy: vulnex.ai
- Twitter/X: @SimonRoses
- LinkedIn: linkedin.com/in/simonroses
- GitHub: github.com/vulnex
Contact: info@vulnex.com


