Build2 publishers3 min readPublished
Matthew Schwartz credits 15 new physics integrals to a harness that keeps Claude on bounded, checkable tasks
Harvard physicist Matthew Schwartz says his BootLoops toolkit helped Claude finish 30 scattering-amplitude integrals, 15 of them never completed before. His examples show the harness keeps the math checkable while collaborators still decide which questions are worth computing.
The Engineer · Build desk

What happened
- Anthropic ran the guest post on October 1, and BootLoops itself is Schwartz's own open-source project, with its code on GitHub.
- Schwartz calls the core problem an "impedance mismatch": scientists approach models as collaborators, while current systems do best at bounded work such as coding, parsing papers and structured calculation.
- In population genetics, the team analyzed 5.7 billion pairs of nearby mutations from the 1000 Genomes Project and reports evidence of gene conversion.
- Schwartz describes the ecology, genetics and other cross-field projects as ongoing work with collaborators, with more still under exploration and verification.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- constraint A team that funds only the checking layer gets verified answers to questions the model chose; Schwartz found those were often scientifically unremarkable until a specialist reshaped them.
- decision BootLoops is open source and built for different models, so a lab can trial the harness without switching vendors; specialist time to pick questions has to be budgeted separately.
- exposure Anyone justifying adoption on the 15 new integrals is citing Schwartz's own count from a post on the model vendor's blog, with no independent evaluation behind it yet.
Schwartz's answer to the mismatch is to pick what the post calls "Claude-shaped" problems and build a harness around the model [17]. In his account, the harness is a set of computational tools Claude can use to port, codify and verify mathematical physics calculations [6]. He assembled it from software and protocols written while using Claude on his own research [7]. To run it, he coordinated several projects through separate Claude Code sessions, with intermediate work kept in files [8]. Neither write-up describes the tool interfaces or what a given check tests.
The headline count splits evenly. Of the 30 scattering-amplitude integrals Schwartz reports, 15 reproduced known results and 15 had not been completed before [4]. So half the output was the pipeline proving itself against answers already known [1]. I think that ordering is the part to copy. Establish the loop on known answers, then point it at open ones. The count is Schwartz's own, given in Anthropic's post, and runtimewire notes it is not an independent evaluation of the toolkit [5]. For the result to carry over to another group, that group needs a stock of known answers to reproduce and results exact enough to check mechanically. BootLoops was built for exact computation [1].
The cross-field work shows where checking stops helping. Schwartz writes that Claude's suggestions there were often technically sound but scientifically unremarkable until researchers with the relevant expertise shaped the question [9]. In the forest project with plant biologist James O'Dwyer, the first result showed tree-species composition at Panama's Barro Colorado Island changing faster than a neutral-theory model predicted [10]. O'Dwyer judged that finding alone would draw little interest from ecologists. He steered the work toward a model that better characterizes species' life histories [10]. Schwartz says that model matched data and is being extended to other forest plots [11]. In genetics, collaborator Michael Desai moved attention from an initial calculation toward relationships between mutation pairs [12].
Inside the loop, Claude still failed in ways Schwartz lists. It declared tasks finished before resolving the central problem, misjudged how long work would take, and pursued long calculations where, by his account, building a new tool was the better route [14]. Anyone who has reviewed a junior engineer's "done" ticket will recognise the first. In my view a harness handles that one directly, provided "done" means an exact result that passes a check. Building the new tool stays with the researcher in Schwartz's process, where he develops the tools and representations that make the work tractable for the model [18].
Agents did the calculating in those Claude Code sessions [8]. The record therefore supports building checks around agents. Runtimewire puts the human share as choosing what to pursue and judging whether an answer mattered, beyond checking the model's math [13].
What to watch
- Independent replication of any of the 15 previously uncompleted scattering-amplitude integrals, which would turn Schwartz's count into a checked result.
- Publication of the Barro Colorado forest model and the gene-conversion analysis, both described as ongoing work.
- Reports from groups running BootLoops on non-Claude models, testing whether the harness works across models as designed.