Science1 publisher2 min readPublished
Natural E. coli mRNAs clump less than synonymous versions of the same genes
Purified bacterial mRNA aggregates in a tube, roughly as physical chemistry predicts. Whitehead researchers report that real E. coli sequences interact less than computer-made synonyms, and read that gap as selection for solubility.
The Scientist · Science desk

What happened
- A team at the Whitehead Institute reported in the Proceedings of the National Academy of Sciences on Sept. 14 that evolution appears to have shaped E. coli mRNA sequences to reduce their interactions with each other.
- When the team purified mRNA out of E. coli and let it behave without its proteins, it aggregated, and sequencing showed the molecules enriched in the clumps matched the simulation's predictions.
- Natural E. coli mRNAs were less prone to interacting with one another than computer-generated alternative sequences that encode the identical proteins.
Compiled by The ScientistSomething wrong?How this is made
Why it matters
- capability RNA drug designers get a variable they can set in the sequence itself, ranking synonymous versions of one construct by how much each tends to associate with other RNA.
- constraint Codon choice has to answer to two requirements at once, so any synonymous optimization aimed at one property is trading against a physical property the authors say evolution has been selecting on.
- decision A program deciding whether to screen synonymous variants for solubility has to make that call on bacterial, computational and protein-free evidence, with no live-cell swap to lean on.
Unwanted RNA interactions can keep mRNAs from being available to make protein, and large RNA aggregates can be toxic to cells [3]. Thousands of mRNA molecules sit crowded into a tiny volume inside a cell, and by their physical properties they ought to be clumping [2]. Purified RNA in a tube does exactly that [1].
Marco Todisco, a postdoctoral researcher in Ankur Jain's lab at the Whitehead Institute, and colleagues turned to the sequences themselves [4]. The genetic code is redundant: most amino acids have more than one codon, so an enormous number of different mRNA sequences encode exactly the same protein [10]. The team used that freedom to build alternative evolutionary histories, generating E. coli mRNAs that differ in their nucleotides and agree in their proteins [11]. Native sequences came out less prone to interacting with one another than the alternatives did, and they tended to fold back onto themselves [12].
The simulation is a physical-chemistry argument about individual molecules at intracellular concentrations, and it predicted that longer mRNAs would be the most likely to join clusters [7]. The purified-RNA experiment tested that prediction, and the molecules enriched in the aggregates had the properties the model flagged [9]. But the assay works by taking away the proteins and other components that normally surround an mRNA [8]. That leaves no way to split a living cell's solubility between codon choice and the components that were washed away. It reports what RNA does on its own.
The sequence comparison is also a contrast between real genes and synonymous variants that never existed [11]. A gap between the two is the signature selection for solubility would leave, and the authors describe it as a previously unrecognized constraint on genetic sequences: DNA has to encode a working protein and also produce an mRNA with physical properties that keep it soluble [14].
The Whitehead account puts the therapeutic implication as a possibility, saying the work could inform RNA drug development by letting designers learn from evolution when they select sequences [15]. The reported experiments stop at E. coli, and none of them tests a therapeutic construct. For an mRNA program, the usable form of this finding would be a score applied across synonymous versions of one construct, ranking how much each version tends to associate with other RNA. Whether such a score predicts anything in a mammalian cell is untested [6].
What to watch
- Whether the same natural-versus-synonymous gap in interaction propensity appears in a eukaryotic transcriptome.
- A live-cell test that swaps a gene's codons for interaction-prone synonyms and then measures aggregation or protein output.
- Any mRNA drug developer reporting a measured stability difference between synonymous versions of one construct.