Science2 distinct publishers3 min readPublished
Protein design has graded its models on whether they can guess the sequence evolution happened to pick. An MIT group says that measures the wrong thing, and reports that leaning less on native sequences improves stability prediction.
The Scientist · Science desk

Compiled by The ScientistSomething wrong?How this is made
Sequence recovery has a denominator problem. Many different amino acid sequences fold into the same structure [3], so a model that proposes something unlike the native protein may have found another workable answer, and a recovery score marks it wrong either way. Amy Keating, who heads MIT's Department of Biology and is senior author on the PNAS paper, frames this as a measurement error rather than a modelling one: the field has asked whether a model can reproduce the sequence "that evolution happened to select", and she says "this isn't the best metric for protein design" [1] [2].
What lead author Foster Birnbaum offers instead are quantities you can score without a native reference: how likely generated sequences are to adopt the target fold, how well the model grasps the sequence-energy landscape, and how well it predicts the effect of a mutation on stability [6]. PottsMPNN is built for the last two. Rather than treating positions one at a time, it carries a pairwise distribution over the physical interactions between all 20 amino acid options at each of two positions, which the team gives as a key reason it models the sequence-energy landscape more accurately than other methods [7]. Multiply it out and that is 400 joint combinations for every pair the model tracks, against 20 numbers per position for a model that scores positions separately [8].
Two training choices push the model off the native answer. Noise, meaning variations added to the structure during training, reduces the tendency to mimic native sequences and widens the range of structures the model can write sequences for [9]. Sets of evolutionarily related sequences teach it that many sequences occupy one fold [10]. Birnbaum concedes the second is, in some ways, still a reliance on natural sequences [11]; the reported direction is that as that reliance falls, structural compatibility and energy prediction improve, including for novel proteins [12].
The thing these accounts do not tell you is how much. Neither the MIT release nor the phys.org version, which reproduces the same text [13], gives an effect size, a comparison figure, or a bench test showing that designed sequences fold and stay folded [14]. The phys.org headline keeps the hedge where it belongs, on "potentially stable" [15]. And the most widely used model in the field was released in 2022 and, by Birnbaum's account, has not been surpassed [5], which is a reminder that a metric argument is how a challenger earns a look rather than how it takes over someone else's pipeline. Birnbaum also says that being able to design any protein at will "enables us to do a potentially scary amount of biological engineering" [16], a sentence worth noting because it comes from the person building the tool.
Ranked by verification strength, evidence, and original report placement.
Amy E. Keating: "For years, the field has measured success by asking whether a model can reproduce the protein sequence that evolution happened to select" and "our work shows that this isn't the best metric for protein design."
Amy E. Keating is head of the Department of Biology, Jay A. Stein (1968) Professor of Biology, professor of biological engineering, and senior author of a paper recently published in PNAS.
In nature, many different amino acid sequences can fold into the same structure, and one sequence can potentially adopt different structures depending on the protein's flexibility or a functional trigger.
Many methods for designing novel proteins are two-step: the structure comes first, then a machine-learning framework generates a repertoire of sequences that could potentially adopt that structure.
Birnbaum, graduate student and lead author: for a completely novel designed structure there would be no native sequence to compare to, and what matters is how likely generated sequences are to fold into the desired structures, how well the model understands the sequence-energy landscape, and how well it predicts the effect of mutations on stability.
PottsMPNN uses a pairwise distribution to capture interactions between amino acids; the ability to account for physical interactions between all 20 possible sequence options at a pair of positions is a key reason it models the sequence-energy landscape more accurately than other methods.
Distinct publishers with included, body-backed reporting in this cluster.
news.mit.edu
1 article · August 27, 2026
phys.org
1 article · August 27, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
science
Odor mixtures get a benchmark: 0.08 median error, 0.57 correlation, 96 pairs of evidence1 distinct publisher
science
A PNAS study says X's feed learns from your arguments, not your likes2 distinct publishers
science
742 species, six continents: what grows back on plowed grassland is not what was lost1 distinct publisher
science
Neanderthal hips had the wide birth canal without the walking penalty1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Peer-reviewed paper behind a single press release, with no figures reported
The cluster's factual base is one institutional press release, republished verbatim, describing a paper in PNAS (DOI given). Peer review and named senior/lead authors lend some weight, and the method description is internally coherent and specific about mechanism. But no quantity, baseline, dataset, ablation, or experimental folding/stability test appears in either account, the incumbent model used as the comparison point is never named, and no independent expert or replicating group is quoted. Descriptive method claims are well grounded; every performance claim is unverifiable from the supplied material.
No adoption signal disclosed
The only dated event in the cluster is the PNAS publication itself. Neither account reports code, weights, a repository, a license, downloads, pipeline integrations, partner labs, or any user of PottsMPNN; the pipeline benefit is stated as a future possibility. A research publication is not an adoption measurement, so this dimension cannot be scored without inventing usage facts.
Aggregator headline and 'design any protein' framing outrun the reported evidence
The originating release is comparatively hedged — it claims better sequence-energy landscape modeling and mutation-effect prediction — but it still asserts improvement without a single number and promises pipeline-level capability that is not demonstrated. Republication then escalates the framing to a headline about an AI model designing potentially stable sequences beyond nature, and the closing quotes stretch to designing 'any protein we want'. Against zero reported benchmarks and zero experimental stability tests, the public framing is moderately overstated; the gap is not larger because the authors themselves flag a real limitation (evolutionary data is still native-sequence reliance) and the underlying work is peer reviewed.
University communications output, republished by an aggregator, with no counterweight
Every word in the cluster originates from the research institution's own news office, quoting the department head who is senior author and the graduate student who is lead author — parties with direct reputational, career, and funding stakes in the framing that the field's dominant benchmark is wrong and that their method fixes it. The republisher adds a traffic-friendly capability headline and bibliographic tags rather than scrutiny. No independent reviewer, competing lab, or critical voice appears, and the promotional framing is not offset by disclosed limitations beyond one authorial caveat.
Clear, self-consistent material but very thin and single-sourced
Confidence in this assessment is moderate. The supplied text is unambiguous, attributed, and internally consistent, and the republication relationship is explicitly stated, so the perspective and incentive readings are secure. Confidence is held down because the substantive scientific question — whether PottsMPNN actually improves sequence-energy landscape modeling and mutation-effect prediction — depends on a PNAS paper that is cited but not contained in the cluster, leaving the evidence and hype-gap scores sensitive to material not supplied.