Build1 distinct publisher3 min readUpdated
A one-day hackathon put agent-designed binders through Adaptyv Bio's wet-lab screen. The durable result is that the designs cleared a shared assay at all, not that any single model won.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
A 34.3% TREM2 binder hit rate is circulating as a Claude result, and it is not one. According to a dev.to analysis of a February 2026 one-day hackathon organised by muni, that number covers the entire agent-designed submission set, screened in the wet lab by Adaptyv Bio, in a campaign where Claude Sonnet 4.6 was one of six agents [1][2][3].
The arithmetic is small and worth stating plainly. The screen evaluated 35 agent-designed binders, of which 12 bound TREM2 [2]. Human participants submitted 65 designs, of which 25 bound, a 38.5% hit rate [5]. That is 100 tested designs and 37 binders across the whole exercise [2], with the human group ahead of the collective agent group by about four percentage points [1]. The dev.to writeup notes that the muni report on the agents-versus-humans experiment is the primary account of the results [4], and that 34.3% sits at the upper end of the mid-teens to mid-30s range cited for Claude-enabled design work [6].
Why the collective framing matters: if 35 designs were spread across six agents, the average is under six designs per agent [3]. No per-model hit rate computed on that denominator carries weight. The source makes the same point in stronger terms, arguing that assessment has to stay model-specific and campaign-specific and that the aggregate agent figure cannot serve as a standalone benchmark for Claude or any other individual system [11].
The affinity data cut against a clean vendor narrative in either direction. One agent-designed binder associated with GPT-5.2 PXDesign was reported at 3.64 nM KD, while a human-designed submission, 13_MRAZS_mosaic, came in at 1.11 nM KD, the strongest reported result in the campaign [7][8]. The best human binder was roughly three times tighter than the best cited agent binder [4]. Both are real binding, not noise at the detection floor.
What is actually new here is procedural, not chemical. Computational protein design has been an active field for years; the consequential step is that agent-generated sequences went into a shared wet-lab screen and returned binders at a non-trivial rate inside a short campaign [10]. One target, one day of design time, one assay format. The source is explicit that this establishes nothing about performance on targets with different structural properties, and nothing about developability, manufacturability, safety, pharmacokinetics, selectivity, or clinical efficacy [9].
For anyone running a design-to-test loop, the operational consequence is triage cost. The bottleneck moves toward deciding which AI-generated candidates deserve follow-up and then pushing them through progressively more demanding assays [12]. Generation is cheap; the queue for the wet lab is not. The same writeup argues that teams running autonomous or semi-autonomous design need explicit controls over model access, design provenance, data handling, human review, and laboratory test criteria [13], which is a reasonable ask when submissions arrive from six systems and the ledger has to survive an audit.
Watch for the per-agent breakdown with design counts attached, a second target with different structural constraints, and whether any of these binders survive a developability screen. Until then the honest headline is that agents can put candidates on the bench, not that they win.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
In a February 2026 one-day hackathon organised by muni, autonomous AI agents including Claude Sonnet 4.6 submitted protein binder designs that were subsequently tested in the wet lab by Adaptyv Bio.
The screen evaluated 35 agent-designed binders, of which 12 bound TREM2, for a 34.3% binder hit rate.
The reported 34.3% figure applies to the collective agent-designed set, not to Claude Sonnet 4.6 alone; Claude was one of six agents in the campaign.
The muni report on the agents-versus-humans TREM2 experiment provides the primary account of the results.
Human participants designed 65 binders, of which 25 bound successfully, a 38.5% hit rate, so the human group outperformed the collective AI-agent group on this measure.
One agent-designed binder associated with GPT-5.2 PXDesign was reported at 3.64 nM KD.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Specific numbers, single second-hand account
The core counts are specific and internally consistent (35/12 agent, 65/25 human, 100 designs and 37 binders overall, with two named KD values), which is more than a vibe-level report. But every figure reaches the cluster through one secondary post that only references muni's report without excerpting or linking assay data, and there is no per-model breakdown, replicate count, binding threshold, or uncertainty estimate. Two of the article's load-bearing interpretive claims -- the mid-teens-to-mid-30s comparison range and the bottleneck shift -- carry no supporting data at all.
One hackathon campaign, no production use
Adoption evidence amounts to a single time-boxed event: one target, one day, one wet-lab screen, 100 total designs. No organisation is reported to have integrated agent-designed binders into an ongoing pipeline, no repeat campaigns or second targets are described, and the source itself states the exercise does not establish cross-target performance. The governance discussion is prescriptive advice, not evidence that any team has adopted such controls.
Mildly overstated, but the source self-corrects
The cluster's headline number is the kind of figure that travels attached to one vendor's model, and the source concedes that the 34.3% is routinely read as Claude's when it belongs to six agents with no per-model breakdown -- that residual mismatch, plus a consultancy call to action riding on the result, keeps the gap positive. It stays small because the article does most of the deflating itself: it names the human group as the better performer and the source of the tightest binder, restricts the claim to one target and one day, and denies that anything here implies faster clinical development.
Vendor-model framing plus a consultancy pitch
The article closes with an explicit commercial solicitation -- a named AI consultancy offering workflow assessment and governance work, with a request-a-consultation call to action -- immediately after arguing that governance and human review are now essential. The headline and lede also attach a named commercial model to a result the body then says cannot be attributed to it, which serves attention even as the text corrects it. No disclosure addresses whether the author or consultancy has any relationship to muni, Adaptyv Bio, or the model vendors.
Coherent single-source account, uncorroborated
Confidence is limited by structure rather than by internal contradiction: one publisher, one article, no primary documents, and a promotional interest in the conclusion. The numbers are consistent and the caveats are unusually candid, which supports moderate trust in the direction of the finding; the specific rates, affinities, and the six-agent composition would need muni's report or Adaptyv's data to firm up, and per-model performance cannot be assessed at all.
product
Anthropic's protein binders got tested by outside labs. The benchmark is still Anthropic's.1 distinct publisher
build
Gemini 3.7 Flash goes GA on one model layer, and that is the actual news1 distinct publisher
build
OpenAI's confirmed NVIDIA footprint is a rack, not a chip; Rubin is still a roadmap1 distinct publisher
build
GPT-5.6 ships as three models, and that makes model choice a deployment decision1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 18, 2026