A Harvard physicist has open-sourced the toolkit he used to turn Claude into a working scientist across 18 fields. On October 1, 2026, Anthropic’s Science blog published a guest post by Matthew Schwartz introducing BootLoops, a harness for “precision quantitative science” that any model can drive. His summary of the state of AI in science is blunt: “Claude and GPT are good at science, but they are not scientists.”
This page is read from Schwartz’s post, the BootLoops repository and bootloops.ai. Everything below about results is the author’s own report.
On this page · 9 sectionsOpen
- On October 1, 2026, Anthropic published a guest post by Harvard physicist Matthew Schwartz introducing BootLoops, an open-source harness for exact, checkable quantitative science.
- BootLoops is a large scientific software package plus working protocols, built so an LLM agent can drive it. It is MIT licensed and works with any model, not only Claude.
- Schwartz reports 36 manuscripts in 18 fields with 19 coauthors over three months, picked from some 400 candidate problems, including 15 Feynman integrals never computed before.
- His central lesson: Claude is good at science but is not a scientist. Its findings were usually technically correct and scientifically dull until a domain expert steered them.
- The results are the author’s own claims. Most projects are still undergoing further verification, and Schwartz has been a visiting researcher at Anthropic. BootLoops is not an Anthropic project.
- The run was compute and token intensive. No cost figure was published.
§ 01What BootLoops is
Schwartz describes it as a harness in the same sense as the coding tools: “much like Claude Code or Claude Science is a harness for Claude, or Codex is a harness for GPT.” The difference is that it does not belong to a model. The project site says it “can be called with Claude, or Gemini or ChatGPT, any version,” and the post adds: “It is also open-source, so it can be used with whatever model you like.”
| Item | Detail |
|---|---|
| Owner | Matthew Schwartz, Harvard; not an Anthropic project |
| Released | October 1, 2026; repository created 15:29 UTC, post published 14:02 UTC |
| License | MIT for code, CC BY 4.0 for documentation |
| Contents | Scientific software ported to one framework, plus protocols, tests and acceptance gates |
| Field codebases | JaCKandJill (phylogenetics), Mixalot (statistics), Terrier (string landscape), Popcorn (population genetics) |
| Model | Built with Claude; usable with any agent |
The name comes from where it started: computing scattering amplitudes, the formulas behind collider predictions, with the S-matrix bootstrap. The site’s rule for when a Feynman diagram counts as done is strict: “a diagram is considered solved if 1) its full functional form is known and 2) a Python script can evaluate it to arbitrary precision on a laptop.” That standard is what makes the work checkable by anyone, expert or not.
§ 02What it produced
The first test was Schwartz’s own field. Claude reproduced results from one of his papers in about 20 minutes, against weeks for his original code, then moved from logarithmic integrals to elliptic ones. “Soon we had 30 integrals BootLooped from end to end, comprising 15 reproductions of known results by this new method and 15 that had never before been computed.”
Then it left physics. The same integrals turned up as Bayesian evidence calculations in population genetics and phylogenetics, and the finite-field methods carried over to evolutionary biology.
| Field | What the post reports |
|---|---|
| Collider physics | 30 Feynman integrals, 15 of them new |
| Ecology | Solved Etienne’s neutral-theory equation; the Barro Colorado Island forest changes 4.5 times faster than neutral theory allows |
| Population genetics | 5.7 billion pairs of nearby mutations analyzed; evidence for gene conversion |
| Economics | 4,452 replication packages from five journals ported to open code, about 30,000 routines; an NBER working paper |
| Linguistics | AccStack, a word-stress database covering 6,072 languages |
| Mathematical physics | An exact solution to Watson’s “final problem”, the 3D random walk with unequal hopping rates |
The scale of the run, in his words: “36 manuscripts in 18 fields with 19 coauthors over three months, out of some 400 candidate problems.”
| Measure | Number |
|---|---|
| Candidate problems | 400 |
| Manuscripts | 36 |
| Coauthors | 19 |
| Fields | 18 |
§ 03The real finding: technically correct is not interesting
The most useful part of the post is the pattern Schwartz kept hitting outside his home field. He found that “Claude was technically correct, but the result was not all that interesting until the expert helped steer us.” The ecology result shows the arc: the plant biologist James O’Dwyer said the first version would likely “be met with a shrug by many ecologists,” then proposed subtracting the neutral prediction and studying what was left. That became the published model.
Schwartz’s name for what the model is good at is a “Claude-shaped” problem: one where a technique from another discipline would solve it outright if anyone in the field knew it existed. The model brings breadth and code. The human brings taste.
§ 04How he runs it
The setup is plain, and it will look familiar to anyone running long agent work. Claude Code sessions run in terminals on Google Cloud machines, one per project, with a master session that coordinates the others, allocates compute and validates results. Intermediate results live in markdown files. Separate sessions write, build the website and act as an adversarial referee.
The failure modes he lists are the ones every team running agents for days meets:
- Declaring victory. “Claude loves to declare victory.” One proof it was proud of held up to “one unproven lemma,” and that lemma was the whole proof.
- No sense of time. Its ETAs were always too long or too short, and it narrated three days of work as a two-year campaign.
- The grind. It would rather run a multiday calculation than build the tool that makes it take minutes. “The model will grind forever if you let it.”
- Context loss. Compaction dropped important context, so he had it consolidate its files and keep the latest plan written down.
§ 05What is verified and what is not
- The harness is real and public. The repository is live under an MIT license with a self-test script; anyone can run it.
- The results are the author’s report. Schwartz says the additional highlights are each “undergoing further exploration and verification.” The economics work has an NBER working paper; most manuscripts are listed on bootloops.ai, not yet in journals.
- The disclosure. Schwartz has been a visiting researcher at Anthropic during the project, and the post appeared on Anthropic’s blog. The post also states: “BootLoops is not an Anthropic project; it is owned and maintained by Matthew Schwartz.”
- The cost. The projects were “compute- and token-intensive.” No dollar figure was published.
§ 06Why it matters
Two things stand out. First, the harness is separate from the model. Anthropic’s nine-loop run last week used Claude Science, a paid platform; BootLoops puts bootstrap machinery from the same field in the open for any agent. Second, Schwartz is careful about what does not change. In his outlook, “humans are still needed for the conceptual part,” and he sees no reason to revisit the scientific method, only to speed it up.
That is the same split we argued for in our essay on self-improving AI: the harness half compounds through use. Schwartz says it plainly about his own: every problem leaves tools behind for the next one. For how harnesses compare on agent work more broadly, see our harness ranking.
§ 07What we are watching for
- Journal acceptance or independent replication of any of the 36 manuscripts.
- Outside contributions to the repository, and runs on models other than Claude.
- A cost figure for a typical BootLoops project.
§ 08Update log
- October 1, 2026: page opened, from Schwartz’s post, the repository and bootloops.ai.
§ 09Sources
- Matthew Schwartz, Claude-shaped science, Anthropic Science blog, October 1, 2026.
- BootLoops on GitHub, MIT license, created October 1, 2026.
- bootloops.ai, the project site with manuscripts and tool pages.
- Anthropic on X, October 1, 2026.
Q1Who built BootLoops?
Matthew Schwartz, a physics professor at Harvard, who owns and maintains it. Anthropic published his guest post but says BootLoops is not an Anthropic project; the GitHub repository says the same.
Q2Is BootLoops free to use?
Yes. The code is MIT licensed and the documentation is CC BY 4.0. Running it costs whatever your model and compute cost; Schwartz calls his projects compute and token intensive and gives no figure.
Q3What is a Claude-shaped problem?
Schwartz’s term for a problem in one field that a known technique from mathematics, physics or computer science would solve outright, if anyone in that field knew the technique existed. Breadth of knowledge plus coding is the model’s edge there.
Q4How does it relate to Claude computing the nine-loop amplitude?
Both come from the S-matrix bootstrap program in particle physics, and both appeared on Anthropic’s Science blog in September and October 2026. The nine-loop run used Claude Science; BootLoops is a separate, open harness any model can drive.
Q5Where does CellCog fit?
We build AI employees for any role, research included. Schwartz’s setup, a coordinating session over per-project sessions with written plans that survive context loss, is the same problem our employees solve with memory and handovers.
