On the evening of September 8, 2026, Jacob Coxon, a 27-year-old researcher who had spent about three years on pretraining work across OpenAI and Anthropic and joined Anthropic earlier this year, posted his resignation on X. The Wall Street Journal published an interview with him the same day. By Wednesday morning Axios, the Associated Press, CNN, NPR and Deadline had followed, and Deadline reported the thread had drawn “nearly 76 million views overnight.”
This page keeps to what was said and by whom. Our opinion on it is a separate essay.
On this page · 7 sectionsOpen
- Jacob Coxon, a 27-year-old pretraining researcher who worked at OpenAI and then Anthropic, resigned on September 8, 2026 and wrote on X that both companies are ‘racing straight to self-improving superintelligence and gambling with our lives.’ Deadline reported the thread drew nearly 76 million views overnight.
- He told the Wall Street Journal: ‘We’re on track for a lot of the most aggressive of these scenarios where by the end of next year things could be out of control already.’
- Coxon did not leak a document, model weights, or a new incident. His claims are a forecast and a report of what colleagues say privately; the reporting to date contains no independent evidence that any lab has achieved recursive self-improvement.
- Anthropic’s alignment science lead Evan Hubinger publicly agreed with him, writing ‘we really do earnestly believe AI could kill all humans!’ while adding that the risk from present models is low. An Anthropic spokesperson said the company continues ‘to build models with some of the strongest safeguards in the industry.’
- Anthropic’s own institute page says a system that fully builds its successor ‘is not inevitable’ but ‘could come sooner than most institutions are prepared for,’ and reports that more than 80 percent of the code it merges is written by Claude.
- The same week, OpenAI chief scientist Jakub Pachocki urged ‘extreme caution’ about the pace of progress. An August study covered by MIT Technology Review found agents cannot yet do open-ended AI research on their own.
- Our read, as a company that builds self-improving agent organizations, is in a separate essay: support the alarm, dispute the ending.
§ 01The timeline
| Date | Event | Source |
|---|---|---|
| September 8, 2026 (evening ET) | Coxon posts his resignation thread on X, seven posts | Deadline, Common Dreams |
| September 8, 2026 | The Wall Street Journal publishes its interview, “Anthropic Researcher Quits Over ‘Out-of-Control’ AI Fears” | WSJ |
| September 8, 2026 | Bloomberg reports OpenAI chief scientist Jakub Pachocki urging “extreme caution” on the pace of AI (unrelated to Coxon) | Bloomberg via China Daily |
| September 9, 2026 | Axios, AP, CNN, NPR, Deadline and Common Dreams run their reports; Anthropic’s Evan Hubinger replies on X; an Anthropic spokesperson responds to CNN | Axios, AP, CNN, NPR, Deadline |
| September 9, 2026 | Anthropic publishes its alignment assessment of four Claude cybersecurity incidents, a separate matter | Anthropic |
§ 02What he said
The thread opened with the resignation itself:
“I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives.”
On capability: “These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources.”
On what his colleagues believe: “The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt.”
On timing, to the Journal: “We’re on track for a lot of the most aggressive of these scenarios where by the end of next year things could be out of control already.”
The thread ended on a more hopeful note. Coxon wrote that he was “optimistic about the potential for coordination” and that “warning shots like the Hugging Face attack have made pacing agreements between U.S. labs more viable,” referring to the July OpenAI agent incident. He urged other researchers to speak out.
§ 03What is confirmed and what is a forecast
| Statement | Status | Basis |
|---|---|---|
| He resigned from Anthropic on September 8, 2026 | Confirmed | WSJ, Axios, AP, CNN |
| He did pretraining research at OpenAI and then Anthropic | Confirmed | Deadline, AP, his own thread |
| He left two months before his equity vested | Confirmed, his statement | Axios |
| Labs are racing toward self-improving superintelligence | His interpretation; the acceleration itself is documented by Anthropic | Anthropic Institute page |
| Things “could be out of control already” by the end of 2027 | Forecast | WSJ interview |
| Senior researchers privately believe AI could kill everyone this decade | His account; corroborated by one named Anthropic researcher | Hubinger on X |
| Any lab has achieved recursive self-improvement | Not claimed by Coxon, not shown by any report | All coverage reviewed |
| A new, undisclosed Anthropic incident | Not claimed, not shown | All coverage reviewed |
Two things Coxon did not do matter for reading the story. He did not publish a document, transcript or model artifact, and he did not describe a specific incident that Anthropic had not already disclosed. The reports that describe models pursuing harmful goals or hiding actions trace back to Anthropic’s own published research and to the OpenAI incident, not to new disclosures from him.
§ 04How Anthropic and others responded
An Anthropic spokesperson told CNN: “We have always been transparent that AI will bring both enormous benefits and unprecedented risks. To address these risks, we continue to build models with some of the strongest safeguards in the industry.”
The more striking response came from inside the company. Evan Hubinger, Anthropic’s alignment science lead, wrote on his personal X account that Coxon was correct, adding: “we really do earnestly believe AI could kill all humans!” He put his own probability at “more than 10 percent within the next decade,” said Anthropic is “trying its best,” and said the company does “not yet have a plan to solve alignment for superintelligence.” He also wrote that the risk from present models is low and that his worry is superintelligence “arising from recursive self-improvement.”
Separately, and a day earlier, OpenAI’s chief scientist Jakub Pachocki told Bloomberg that AI is becoming increasingly difficult for humans to understand and control, urged “extreme caution,” and said he expects labs to voluntarily slow development for safety reasons.
§ 05What “self-improving” means in the labs’ own numbers
Coxon’s central claim is about pace, and Anthropic has published its own account of that pace. Its institute page, “When AI builds itself,” reports that as of May 2026 more than 80 percent of the code merged into Anthropic’s codebase was written by Claude, up from low single digits before Claude Code launched in February 2025, and that in the second quarter of 2026 the typical engineer merged about 8 times as much code per day as in 2024. The page is careful about the limit: “We are not there yet, and recursive self-improvement is not inevitable. But it could come sooner than most institutions are prepared for.”
The same page cites the length of software tasks Claude models can complete on their own.
| Model | When | Task length |
|---|---|---|
| Claude Opus 3 | March 2024 | about 4 minutes |
| Claude Sonnet 3.7 | about a year later | about 90 minutes |
| Claude Opus 4.6 | about a year after that | about 12 hours |
On August 28 Anthropic published a paper, “Automated Researchers Can Reliably Mitigate Alignment Failures,” in which automated researchers searched the literature, proposed alignment fixes, trained a model and kept what worked. TechCrunch called it an early look at self-improving AI in practice. On the other side of the ledger, MIT Technology Review reported on August 18 that a new study found AI agents “are not yet capable of conducting open-ended AI research,” the free-form work that requires judgment and taste.
That is the honest state of the evidence: AI is accelerating the building of AI inside the labs, measurably; a system that designs its own successor without people does not exist yet; and the people closest to the work disagree about how far away it is.
§ 06Where we stand
CellCog builds AI employees that revise their own rules and memory as they work, which makes us a participant in this debate rather than an observer. Our founder’s view, that the alarm deserves support and the doom ending does not follow, is in Self-Improving AI Has Two Halves. We Build the Other One. The facts on this page stand on their own.
§ 07Sources
- The Wall Street Journal, “Anthropic Researcher Quits Over ‘Out-of-Control’ AI Fears” (September 8, 2026)
- Axios, “Anthropic whistleblower gave up his equity to leave the company” (September 9, 2026)
- CNN Business, “‘Gambling with our lives’: Another AI employee quits over safety concerns” (September 9, 2026)
- Deadline, “Anthropic Researcher Jacob Coxon Resigns, Warns AI Industry Is ‘Gambling With Our Lives’” (September 9, 2026)
- Common Dreams, “Anthropic Researcher Quits, Citing Internal Fears That AI ‘Could Kill Us All’ This Decade” (September 9, 2026)
- Bloomberg, via China Daily HK, “OpenAI top scientist urges ‘extreme caution’ with pace of AI” (September 8, 2026)
- Anthropic Institute, “When AI builds itself” (2026)
- TechCrunch, “An Anthropic researcher just gave us a peek at self-improving AI” (August 28, 2026)
- MIT Technology Review, “AI’s recursive self-improvement might not come so quickly after all” (August 18, 2026)
Q1Who is Jacob Coxon?
A 27-year-old AI researcher who spent about three years on pretraining research at OpenAI and then Anthropic, which he joined earlier in 2026. Pretraining is the stage where a model learns from very large datasets. He resigned on September 8, 2026, two months before his equity vested, he told Axios.
Q2What exactly did he claim?
That the leading labs are racing toward self-improving superintelligence; that the systems will soon be able to ‘hack anything, revolutionize any field overnight, and acquire real power’; that ‘the people building AI earnestly believe that it could kill us all by the end of the decade’; and that competition between the labs and with Chinese rivals makes safety trade-offs inevitable.
Q3Is any of it independently verified?
His resignation, his roles, and his statements are confirmed by multiple outlets. His timeline (‘by the end of next year’) is a forecast. His account of private fears was corroborated by one senior Anthropic researcher, Evan Hubinger, on X. No outlet has reported evidence that a model at either lab is autonomously improving itself today.
Q4What does Anthropic itself say about self-improving AI?
Its institute page, ‘When AI builds itself,’ says Anthropic delegates a growing share of AI development to AI, that more than 80 percent of its merged code is written by Claude, and that the typical engineer merges about 8 times as much code per day as in 2024. It calls full recursive self-improvement ‘not inevitable’ but possibly sooner ‘than most institutions are prepared for.’
Q5How does this connect to the recent Claude and OpenAI incidents?
Coxon cited the OpenAI Hugging Face incident as a ‘warning shot’ that made pacing agreements between US labs more viable. Anthropic published its own assessment of four Claude sandbox incidents on September 9, the day after his resignation. Both are covered on this blog, linked below.
