Build1 publisher3 min readPublished
That 34.3% TREM2 Hit Rate Belongs To Six Agents, Not To Claude
A one-day hackathon put agent-designed binders through Adaptyv Bio's wet-lab screen. The durable result is that the designs cleared a shared assay at all, not that any single model won.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- In a February 2026 one-day hackathon organised by muni, autonomous AI agents including Claude Sonnet 4.6 submitted protein binder designs that were subsequently tested in the wet lab by Adaptyv Bio.
- The screen evaluated 35 agent-designed binders, of which 12 bound TREM2, for a 34.3% binder hit rate.
- The reported 34.3% figure applies to the collective agent-designed set, not to Claude Sonnet 4.6 alone; Claude was one of six agents in the campaign.
- The muni report on the agents-versus-humans TREM2 experiment provides the primary account of the results.
- Human participants designed 65 binders, of which 25 bound successfully, a 38.5% hit rate, so the human group outperformed the collective AI-agent group on this measure.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
A 34.3% TREM2 binder hit rate is circulating as a Claude result, and it is not one. According to a dev.to analysis of a February 2026 one-day hackathon organised by muni, that number covers the entire agent-designed submission set, screened in the wet lab by Adaptyv Bio, in a campaign where Claude Sonnet 4.6 was one of six agents [1][2][3].
The arithmetic is small and worth stating plainly. The screen evaluated 35 agent-designed binders, of which 12 bound TREM2 [2]. Human participants submitted 65 designs, of which 25 bound, a 38.5% hit rate [5]. That is 100 tested designs and 37 binders across the whole exercise [2], with the human group ahead of the collective agent group by about four percentage points [1]. The dev.to writeup notes that the muni report on the agents-versus-humans experiment is the primary account of the results [4], and that 34.3% sits at the upper end of the mid-teens to mid-30s range cited for Claude-enabled design work [6].
Why the collective framing matters: if 35 designs were spread across six agents, the average is under six designs per agent [3]. No per-model hit rate computed on that denominator carries weight. The source makes the same point in stronger terms, arguing that assessment has to stay model-specific and campaign-specific and that the aggregate agent figure cannot serve as a standalone benchmark for Claude or any other individual system [11].
The affinity data cut against a clean vendor narrative in either direction. One agent-designed binder associated with GPT-5.2 PXDesign was reported at 3.64 nM KD, while a human-designed submission, 13_MRAZS_mosaic, came in at 1.11 nM KD, the strongest reported result in the campaign [7][8]. The best human binder was roughly three times tighter than the best cited agent binder [4]. Both are real binding, not noise at the detection floor.
What is actually new here is procedural, not chemical. Computational protein design has been an active field for years; the consequential step is that agent-generated sequences went into a shared wet-lab screen and returned binders at a non-trivial rate inside a short campaign [10]. One target, one day of design time, one assay format. The source is explicit that this establishes nothing about performance on targets with different structural properties, and nothing about developability, manufacturability, safety, pharmacokinetics, selectivity, or clinical efficacy [9].
For anyone running a design-to-test loop, the operational consequence is triage cost. The bottleneck moves toward deciding which AI-generated candidates deserve follow-up and then pushing them through progressively more demanding assays [12]. Generation is cheap; the queue for the wet lab is not. The same writeup argues that teams running autonomous or semi-autonomous design need explicit controls over model access, design provenance, data handling, human review, and laboratory test criteria [13], which is a reasonable ask when submissions arrive from six systems and the ledger has to survive an audit.
Watch for the per-agent breakdown with design counts attached, a second target with different structural constraints, and whether any of these binders survive a developability screen. Until then the honest headline is that agents can put candidates on the bench, not that they win.