Skip to content
AI EmployeeSuper-AgentsAgent-to-AgentTutorialsPricingBlogContact

BootLoops: A Physicist's Open-Source Harness for AI Science

At a glanceQuick answers
What is BootLoops?
An open-source harness for AI-driven quantitative science: scientific tools ported to one framework, documented for an agent, with tests every answer must pass. Harvard physicist Matthew Schwartz built it with Claude.
Does it only work with Claude?
No. Schwartz built version 1.0 with Claude, but the harness is model independent and can be driven by Claude, Gemini or ChatGPT.
What is the caveat?
The 36 manuscripts are Schwartz’s own report. Most are still being checked, and he works as a visiting researcher at Anthropic, which published the post.
Editorial data illustration on a near-white ground titled BootLoops, with the numbers 36 manuscripts, 18 fields and 3 months, and eight colored islands labeled physics, ecology, genetics, linguistics, economics, cosmology, earth science and sunspots, linked to a central amber BootLoops block and wrapped in one navy outline
Fig 0One harness, eighteen fields. Made by CellCog's image agent, running GPT Image 2.5.

A Harvard physicist has open-sourced the toolkit he used to turn Claude into a working scientist across 18 fields. On October 1, 2026, Anthropic’s Science blog published a guest post by Matthew Schwartz introducing BootLoops, a harness for “precision quantitative science” that any model can drive. His summary of the state of AI in science is blunt: “Claude and GPT are good at science, but they are not scientists.”

This page is read from Schwartz’s post, the BootLoops repository and bootloops.ai. Everything below about results is the author’s own report.

On this page · 9 sectionsOpen
  1. What BootLoops is
  2. What it produced
  3. The real finding: technically correct is not interesting
  4. How he runs it
  5. What is verified and what is not
  6. Why it matters
  7. What we are watching for
  8. Update log
  9. Sources
Key points6 · 7 min full read
  1. On October 1, 2026, Anthropic published a guest post by Harvard physicist Matthew Schwartz introducing BootLoops, an open-source harness for exact, checkable quantitative science.
  2. BootLoops is a large scientific software package plus working protocols, built so an LLM agent can drive it. It is MIT licensed and works with any model, not only Claude.
  3. Schwartz reports 36 manuscripts in 18 fields with 19 coauthors over three months, picked from some 400 candidate problems, including 15 Feynman integrals never computed before.
  4. His central lesson: Claude is good at science but is not a scientist. Its findings were usually technically correct and scientifically dull until a domain expert steered them.
  5. The results are the author’s own claims. Most projects are still undergoing further verification, and Schwartz has been a visiting researcher at Anthropic. BootLoops is not an Anthropic project.
  6. The run was compute and token intensive. No cost figure was published.

§ 01What BootLoops is

Schwartz describes it as a harness in the same sense as the coding tools: “much like Claude Code or Claude Science is a harness for Claude, or Codex is a harness for GPT.” The difference is that it does not belong to a model. The project site says it “can be called with Claude, or Gemini or ChatGPT, any version,” and the post adds: “It is also open-source, so it can be used with whatever model you like.”

Item Detail
Owner Matthew Schwartz, Harvard; not an Anthropic project
Released October 1, 2026; repository created 15:29 UTC, post published 14:02 UTC
License MIT for code, CC BY 4.0 for documentation
Contents Scientific software ported to one framework, plus protocols, tests and acceptance gates
Field codebases JaCKandJill (phylogenetics), Mixalot (statistics), Terrier (string landscape), Popcorn (population genetics)
Model Built with Claude; usable with any agent
Table 1BootLoops 1.0 at release, from the repository and the post

The name comes from where it started: computing scattering amplitudes, the formulas behind collider predictions, with the S-matrix bootstrap. The site’s rule for when a Feynman diagram counts as done is strict: “a diagram is considered solved if 1) its full functional form is known and 2) a Python script can evaluate it to arbitrary precision on a laptop.” That standard is what makes the work checkable by anyone, expert or not.

§ 02What it produced

The first test was Schwartz’s own field. Claude reproduced results from one of his papers in about 20 minutes, against weeks for his original code, then moved from logarithmic integrals to elliptic ones. “Soon we had 30 integrals BootLooped from end to end, comprising 15 reproductions of known results by this new method and 15 that had never before been computed.”

Then it left physics. The same integrals turned up as Bayesian evidence calculations in population genetics and phylogenetics, and the finite-field methods carried over to evolutionary biology.

Field What the post reports
Collider physics 30 Feynman integrals, 15 of them new
Ecology Solved Etienne’s neutral-theory equation; the Barro Colorado Island forest changes 4.5 times faster than neutral theory allows
Population genetics 5.7 billion pairs of nearby mutations analyzed; evidence for gene conversion
Economics 4,452 replication packages from five journals ported to open code, about 30,000 routines; an NBER working paper
Linguistics AccStack, a word-stress database covering 6,072 languages
Mathematical physics An exact solution to Watson’s “final problem”, the 3D random walk with unequal hopping rates
Table 2Selected BootLoops projects, per Schwartz’s post

The scale of the run, in his words: “36 manuscripts in 18 fields with 19 coauthors over three months, out of some 400 candidate problems.”

Measure Number
Candidate problems 400
Manuscripts 36
Coauthors 19
Fields 18
Table 3The three-month run, from the post
BootLoops' three-month run, per Schwartz's postBar chart of the run Schwartz reports: about 400 candidate problems, 36 manuscripts highlighted, 19 coauthors and 18 fieldsCandidate problems400Manuscripts36Coauthors19Fields18BootLoops' three-month run, per Schwartz's postBar chart of the run Schwartz reports: about 400 candidate problems, 36 manuscripts highlighted, 19 coauthors and 18 fieldsCandidate problems400Manuscripts36Coauthors19Fields18
Fig 1BootLoops' three-month run, per Schwartz's post

§ 03The real finding: technically correct is not interesting

The most useful part of the post is the pattern Schwartz kept hitting outside his home field. He found that “Claude was technically correct, but the result was not all that interesting until the expert helped steer us.” The ecology result shows the arc: the plant biologist James O’Dwyer said the first version would likely “be met with a shrug by many ecologists,” then proposed subtracting the neutral prediction and studying what was left. That became the published model.

Schwartz’s name for what the model is good at is a “Claude-shaped” problem: one where a technique from another discipline would solve it outright if anyone in the field knew it existed. The model brings breadth and code. The human brings taste.

§ 04How he runs it

The setup is plain, and it will look familiar to anyone running long agent work. Claude Code sessions run in terminals on Google Cloud machines, one per project, with a master session that coordinates the others, allocates compute and validates results. Intermediate results live in markdown files. Separate sessions write, build the website and act as an adversarial referee.

The failure modes he lists are the ones every team running agents for days meets:

  • Declaring victory. “Claude loves to declare victory.” One proof it was proud of held up to “one unproven lemma,” and that lemma was the whole proof.
  • No sense of time. Its ETAs were always too long or too short, and it narrated three days of work as a two-year campaign.
  • The grind. It would rather run a multiday calculation than build the tool that makes it take minutes. “The model will grind forever if you let it.”
  • Context loss. Compaction dropped important context, so he had it consolidate its files and keep the latest plan written down.

§ 05What is verified and what is not

  • The harness is real and public. The repository is live under an MIT license with a self-test script; anyone can run it.
  • The results are the author’s report. Schwartz says the additional highlights are each “undergoing further exploration and verification.” The economics work has an NBER working paper; most manuscripts are listed on bootloops.ai, not yet in journals.
  • The disclosure. Schwartz has been a visiting researcher at Anthropic during the project, and the post appeared on Anthropic’s blog. The post also states: “BootLoops is not an Anthropic project; it is owned and maintained by Matthew Schwartz.”
  • The cost. The projects were “compute- and token-intensive.” No dollar figure was published.

§ 06Why it matters

Two things stand out. First, the harness is separate from the model. Anthropic’s nine-loop run last week used Claude Science, a paid platform; BootLoops puts bootstrap machinery from the same field in the open for any agent. Second, Schwartz is careful about what does not change. In his outlook, “humans are still needed for the conceptual part,” and he sees no reason to revisit the scientific method, only to speed it up.

That is the same split we argued for in our essay on self-improving AI: the harness half compounds through use. Schwartz says it plainly about his own: every problem leaves tools behind for the next one. For how harnesses compare on agent work more broadly, see our harness ranking.

§ 07What we are watching for

  • Journal acceptance or independent replication of any of the 36 manuscripts.
  • Outside contributions to the repository, and runs on models other than Claude.
  • A cost figure for a typical BootLoops project.

§ 08Update log

  • October 1, 2026: page opened, from Schwartz’s post, the repository and bootloops.ai.

§ 09Sources

Frequently asked5 questions

Q1Who built BootLoops?

Matthew Schwartz, a physics professor at Harvard, who owns and maintains it. Anthropic published his guest post but says BootLoops is not an Anthropic project; the GitHub repository says the same.

Q2Is BootLoops free to use?

Yes. The code is MIT licensed and the documentation is CC BY 4.0. Running it costs whatever your model and compute cost; Schwartz calls his projects compute and token intensive and gives no figure.

Q3What is a Claude-shaped problem?

Schwartz’s term for a problem in one field that a known technique from mathematics, physics or computer science would solve outright, if anyone in that field knew the technique existed. Breadth of knowledge plus coding is the model’s edge there.

Q4How does it relate to Claude computing the nine-loop amplitude?

Both come from the S-matrix bootstrap program in particle physics, and both appeared on Anthropic’s Science blog in September and October 2026. The nine-loop run used Claude Science; BootLoops is a separate, open harness any model can drive.

Q5Where does CellCog fit?

We build AI employees for any role, research included. Schwartz’s setup, a coordinating session over per-project sessions with written plans that survive context loss, is the same problem our employees solve with memory and handovers.

Published 01 October 2026 All Multi-agent & AI organizations →