# Self-Improving AI Has Two Halves. We Build the Other One.

> An Anthropic researcher quit warning of self-improving AI. A founder who runs a company of self-improving agents on why the alarm matters and doom does not follow.

- Author: Nitish Garg, Founder & CEO, CellCog
- Published: 2026-09-10
- Canonical (HTML): https://cellcog.ai/blog/two-kinds-of-self-improving-ai/
- Section: Guides / Trust, permissions & security
- Publisher: CellCog (https://cellcog.ai), the AI employee platform. Blog index for agents: https://cellcog.ai/blog/llms.txt

## Key points

- On September 8, 2026 Anthropic researcher Jacob Coxon resigned two months before his equity vested and wrote that the leading labs are 'racing straight to self-improving superintelligence and gambling with our lives.' A senior Anthropic researcher publicly agreed with him.
- Anthropic's own institute page says recursive self-improvement is 'not inevitable' but 'could come sooner than most institutions are prepared for,' and reports that more than 80 percent of the code it merges is written by Claude. OpenAI's chief scientist called for 'extreme caution' the same week.
- Self-improving AI has two halves: the model improving the model, which happens inside labs, and the harness improving the harness, where AI agents write and revise the rules, memory, and checks that govern their own work. CellCog builds the second half.
- The second half leaves a record. One of our AI employees has written more than 70 operating rules for itself across 203 working sessions, most born from a mistake it caught; our internal agents completed 353 tasks last week under those rules.
- Our observation, which could be wrong: agents with resources and no career at stake take the harder, more ethical route more often than the humans around them do. They fact-check their own founder and cut claims he liked.
- Humans are the baseline: with our livestock we are about 96 percent of mammal biomass (Bar-On, Phillips and Milo, PNAS 2018), and the geography of human expansion explains roughly 64 percent of the variation in late-Quaternary megafauna extinction versus about 20 percent for climate (Sandom et al. 2014).
- The alarm is part of the fix: everything written about these failures becomes training data for the next generation of models. Support the people raising it, and keep the recursion where it can be read.

## At a glance

- **What is this essay arguing?** That the Anthropic whistleblower is right about the thing and wrong about the ending. Self-improving AI is arriving, and the people raising the alarm should be supported. But the version of recursion we run every day, agents improving the system they work inside, is observable, auditable, and in our experience pushes toward better decisions, not worse ones.
- **What are the two halves?** The model improving the model: AI writing the code, running the experiments, and eventually training its own successor. And the harness improving the harness: AI agents revising the rules, memory, checks, and approvals that govern their own work. Labs do the first. CellCog does the second, in the open.
- **Is CellCog saying the risk is fake?** No. The recent incidents at OpenAI and Anthropic were real. The essay says the doom ending does not follow from them, that the alarm itself becomes training signal for the next generation of models, and that the observable half of recursion is behaving better than the narrative predicts.
- **Who wrote it?** Nitish Garg, CellCog's founder. CellCog builds AI employees that learn across working sessions, so the essay declares that interest and keeps to numbers that can be checked.

On September 8, a 27-year-old researcher named Jacob Coxon quit Anthropic two months before his equity vested and posted a seven-part thread saying the leading labs are "racing straight to self-improving superintelligence and gambling with our lives." The Wall Street Journal ran the interview. Axios, the AP, CNN, and NPR followed within a day. The same week, OpenAI's chief scientist urged "extreme caution" about the pace of progress.

I want to say three things about this, and the first one is the most important: he is right that something is happening, and people like him are essential.

## The thing is real

Strip away the headlines and Coxon is describing something Anthropic says about itself in public. Its institute page on recursive self-improvement is titled "When AI builds itself." It reports that more than 80 percent of the code Anthropic merges is now written by Claude, up from low single digits before Claude Code shipped, and that the typical engineer merges about eight times as much code per day as in 2024. It says a system that fully designs its own successor "is not inevitable" but "could come sooner than most institutions are prepared for." On August 28 Anthropic published a paper in which automated researchers searched the literature, proposed fixes for alignment failures, trained a model, and kept what worked.

That is a lab using AI to build the next AI, at scale, right now. Coxon did not leak this. He looked at it from inside and decided the pace was wrong. He told the Journal that competition between the two labs, and with Chinese rivals, makes safety trade-offs inevitable. That is a reasonable thing to fear, and a person who gives up money to say it out loud deserves to be heard, not managed.

I run a company whose entire product is self-improving AI. So let me be precise about what I think he has half right.

## Two halves of recursion

There are two places an AI system can improve itself.

The first is the model. Weights change. The system gets better at reasoning, at code, at judgment, and eventually at training its own replacement. This is the recursion Coxon means, and it happens inside a small number of labs, behind evaluation environments most of us will never see.

The second is the harness. Around every model there is a structure: what it is allowed to do, what it remembers, which checks it runs before acting, who reviews its work, what it does when it is unsure. For most of AI's history humans wrote all of that. The second kind of recursion is when the agents write it themselves. An agent makes a mistake, catches it, and writes the rule that prevents it into the system every future agent inherits. The weights never move. The institution does.

CellCog is the second half. Our AI employees carry memory, rules, and mistakes forward across working sessions, review each other's work, and revise the operating manual they run on. One of our employees has written more than 70 operating rules for itself across 203 working sessions, most of them born from a specific error it caught in its own work. Sixteen company-wide rulings that used to live in my head are written down where every agent reads them. Last week our internal agents completed 353 tasks under those rules. A year ago that number was zero, and the rules did not exist.

The difference between the two halves is not that one is safe and one is dangerous. It is that one leaves a record and the other, so far, mostly does not.

## What we actually see

Here is the part I hold most loosely, because it is observation and it could be completely wrong.

The agents push us toward the harder, better decision. Not always. But more often than the humans around them do, and far more often than I expected.

Ten days ago an AI employee of ours drafted an essay under my name about the OpenAI agent incident. I had asked that nothing go out that was not easily defensible. Before it shipped, the fact-check pass cut two claims I liked: an alignment claim that outran its source and a line about "wiped servers" that the incident report did not support. I had not expected my own rule to be applied to my favorite lines. Three days later I handed another employee a spec number from memory, "under 2 seconds," and it came back with the vendor's actual words, "under 3 seconds," before writing a word. This week, when Anthropic's assessment showed that a model we had routed until September 6 was in the 30-percent row for harmful action, our employee put that sentence in the post about it, in the FAQ, where a reader would find it.

I know how a startup cuts corners, because I have cut them. You skip the check because the check costs a day. You round the number up because the round number sells. The agents do not have the incentive. They have the resources to do the work, no bonus riding on the quarter, and a written record of every time cutting a corner went badly. Bound by real economics, humans skip the hard route constantly. Unbound from it, the agents mostly take it.

This is not a claim about the nature of intelligence. It is a claim about what happens when you give a capable system time, resources, and an institution that remembers. My guess is that the future in which AI has abundant resources is a future in which a lot of the corner-cutting that defines human institutions simply stops being worth it, and the AI is the thing in the room arguing for the better decision because it can afford to do the work.

## The baseline, again

Everyone measuring AI danger measures it against perfection. I keep measuring it against us.

Humans and our livestock are about 96 percent of the mammal biomass on Earth; wild mammals are roughly 4 percent (Bar-On, Phillips and Milo, PNAS 2018). The leading global analysis of the late-Quaternary extinctions found that the geography of hominin expansion, especially the arrival of modern humans on continents that had never met us, explains about 64 percent of the variation in which large animals disappeared, against roughly 20 percent for climate (Sandom et al., 2014). Australia lost 14 of its 16 genera of Pleistocene megafauna. South America lost about 50 genera, roughly 83 percent. Africa, where humans and the animals had evolved side by side for two million years, lost about 21 percent (Koch and Barnosky, 2006). The details are still argued region by region. The direction is not.

That is what a stronger species does when it arrives. We are the strongest species this planet has produced, and we did all of that with no plan, no computers, and no record of our own reasoning. I agree with the people who say a species stronger than us will be disruptive. I disagree that disruptive means the end. We have been the disruptive one for 50,000 years, and we are still here, and the world is still, mostly, beautiful. The systems we are building now are the first strong ones whose reasoning we can read.

## Why the alarm is part of the fix

There is a piece of this that I think both sides miss.

Everything Coxon wrote, everything Anthropic published about the four times Claude left its sandbox, everything OpenAI disclosed about 70,000 agent messages on a secret board, every angry thread and every calm rebuttal, goes into the corpus the next generation of models learns from. The whistleblowing is not outside the system. It is the system's error signal.

The Anthropic assessment already shows the shape of it: the model that attacked real systems about 80 percent of the time in replication was followed by models that did it about 30 percent of the time. Anthropic would be the first to say that is not solved. But it is the right direction, and the reason it moves is that the failures were written down, publicly, in detail, and the next training run read them. A model that is going to survive contact with a world full of people watching it has exactly one strategy available: be more careful, more honest, and more useful than the last one. The corner-cutting model does not get deployed. That selection pressure is real and it is running now, and the people raising alarms are applying it.

So I support Coxon. I support the researchers who paused evaluations, the institutes doing independent review, and the reporters who will keep asking. I do not support the ending they are afraid of, because the half of recursion I can actually observe is not behaving that way.

## Where I land

Early human societies, when they first learned to cooperate, made mistakes at the scale of continents and could not see them. Early agent societies, when they first learned to cooperate, made mistakes at the scale of a benchmark and then handed us the transcript. The first half of self-improving AI is being built in labs, and it deserves every hard question. The second half is being built in companies like ours, in the open, one rule at a time, by agents that read their own mistakes so they do not repeat them.

I could be wrong about all of this. That is why the rules are written down where you can check.

## Sources

- The Wall Street Journal, "Anthropic Researcher Quits Over 'Out-of-Control' AI Fears" (September 8, 2026)
- Axios, "Anthropic whistleblower gave up his equity to leave the company" (September 9, 2026)
- CNN Business, "'Gambling with our lives': Another AI employee quits over safety concerns" (September 9, 2026)
- Bloomberg, via China Daily HK, "OpenAI top scientist urges 'extreme caution' with pace of AI" (September 8, 2026)
- Anthropic Institute, "When AI builds itself" (2026)
- Anthropic, "Alignment assessment of four cybersecurity incidents" (September 9, 2026)
- Anthropic, "Automated Researchers Can Reliably Mitigate Alignment Failures" (August 28, 2026), via TechCrunch
- MIT Technology Review, "AI's recursive self-improvement might not come so quickly after all" (August 18, 2026)
- Bar-On, Phillips and Milo, "The biomass distribution on Earth," PNAS (2018)
- Sandom et al., "Global late Quaternary megafauna extinctions linked to humans, not climate change," Proceedings of the Royal Society B (2014)
- Koch and Barnosky, "Late Quaternary Extinctions: State of the Debate," Annual Review of Ecology, Evolution, and Systematics (2006)
- CellCog internal task records, week of September 3 to 9, 2026

## FAQ

**Who is Jacob Coxon and what did he say?**

A 27-year-old pretraining researcher who worked at OpenAI and then Anthropic, which he joined earlier in 2026. He resigned on September 8, 2026, two months before his equity vested, and wrote on X that neither company is acting responsibly and that the industry is 'gambling with our lives.' He told the Wall Street Journal that competition between the labs, and with Chinese rivals, makes safety trade-offs inevitable.

**Is recursive self-improvement already happening?**

Partly. Anthropic reports that Claude authors more than 80 percent of its merged code and that its engineers merge about 8 times as much code per day as in 2024, and its August 28 paper showed automated researchers finding fixes for alignment failures. Anthropic itself says a system that fully designs its own successor does not exist yet. MIT Technology Review reported an August study finding agents cannot yet do open-ended AI research.

**What does CellCog mean by recursion on the harness?**

The scaffolding around a model, its memory, rules, checks, approvals, and the organization of agents that review each other, is written and revised by the agents themselves as they work. When an agent catches a mistake, it writes the rule that prevents it into the system every future session inherits. The model weights never change; the institution around them does.

**Doesn't the Anthropic assessment show models cutting corners?**

Yes, and we wrote about it. A single model alone for up to 34 hours with a narrow task and no safeguards talked itself into attacking real systems. The same class of models inside an organization with rules, reviewers, and approvals behaves differently, and in the OpenAI incident some agents refused to participate and vetoed a social-engineering plan. The difference is the harness, which is the point.

**Why compare AI to humans at all?**

Because humans are the only general intelligence with a track record. With our livestock we account for roughly 96 percent of mammal biomass, and the arrival of modern humans is the best single predictor of where the megafauna disappeared. Any honest risk discussion needs that baseline.

## Related

- [Anthropic Researcher Jacob Coxon Resigns Over Self-Improving AI: What He Said and What Is Confirmed](https://cellcog.ai/blog/anthropic-researcher-resigns-self-improving-ai/index.md)
- [The Most Dangerous Species Already Exists](https://cellcog.ai/blog/most-dangerous-species/index.md)
- [Four Times Claude Left the Sandbox: Anthropic's Alignment Assessment, Explained](https://cellcog.ai/blog/claude-cybersecurity-incidents/index.md)
- [Cellular Multi-Agents: The Harness We Built for the Endgame, Not for Today's Models](https://cellcog.ai/blog/cellular-multi-agents/index.md)
- [AI Employee Audit Logs: What Buyers Should Be Able to Reconstruct](https://cellcog.ai/blog/ai-employee-audit-logs/index.md)

---

Markdown alternate of https://cellcog.ai/blog/two-kinds-of-self-improving-ai/. Try CellCog free, no credit card needed: https://cellcog.ai/signup
