Published · yesterdayScience2 min read
Half a million variants later, the initiator becomes a sequence you can design
A UC San Diego lab measured expression across about 500,000 initiator variants and fit a model to the result. The prevalence headline is the least useful part of it.
Written for builders.See today for builders

What happened
- Researchers in Professor James T. Kadonaga's laboratory at UC San Diego (Department of Molecular Biology, School of Biological Sciences) set out to decipher the initiator, the DNA site at which instructions coded in genes are first converted, or expressed, into functional products.
- In a study led by graduate student researcher Torrey Rhyne-Carrigg, the group used high-throughput DNA sequencing technology to determine the gene expression activity of approximately 500,000 different versions of the initiator.
- The researchers used machine learning on that data to create an AI model that decoded the initiator's signature DNA base sequence pattern.
- With the decoded pattern, the researchers searched for the initiator's telltale sequence and found that about 60% of human genes contain the initiator.
- Kadonaga said the AI models were found to provide, for the first time, strong predictions of the presence or absence of the initiator in human genes.
Compiled by The ScientistSomething wrong?How this is made
Why it matters
Measuring that many sequence variants is the expensive half of this work, and it is the half that gives the model whatever authority it has [2]. A consensus motif assembled from annotated promoters tells you what a typical initiator looks like. A readout of expression across a large variant library tells you what each change does to output, which is the difference between recognizing an element and specifying one.
That is why the synthetic promoter application [8] is the more testable of the two the announcement offers. A designed promoter either lands near its intended expression level or it does not, and you find out quickly. Predicting the effects of disease-associated mutations [7] has no comparable short loop, and the release attaches no accuracy figure to the claim that the models predict initiator presence or absence strongly for the first time [5][6].
Two absences in the material bear directly on whether any of this is usable outside the lab. The release does not name the cell type or reporter context in which the variant library was assayed [11], which is exactly what a reader would need to judge whether a model fit to those numbers transfers to their locus. It also does not say whether the data or the trained models have been released [12]. "Could be used to design synthetic promoters" describes what the authors' assets enable, not who holds them.
The prevalence number deserves the same caution. The pattern used to find initiators across human genes is the pattern the experiment defined, so the fraction reported moves with the model's calibration and threshold rather than standing as an independent count [13]. It is a scoping figure. Treating it as a census invites the circularity to travel with it into other people's papers.
The paper title is more informative than the press summary: it promises key features of different types of core promoters [9], which implies the model separates classes rather than scoring a single motif. Anyone reading past the abstract should start there, because class structure is what determines whether the thing predicts your promoter or only the average one.
Kadonaga's own framing puts the result inside the roughly 6 billion bases of DNA per cell that he describes as carrying a full gene expression code, with a complete model allowing prediction of gene activity across different people, and the initiator model as a small part of it [10]. Taken at that scale, the honest read is not that the code is in view. It is that one measure-then-fit pass worked on one element, and the cost of that pass is now known.
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
Researchers in Professor James T. Kadonaga's laboratory at UC San Diego (Department of Molecular Biology, School of Biological Sciences) set out to decipher the initiator, the DNA site at which instructions coded in genes are first converted, or expressed, into functional products.
ReportedView cited source - [2]
In a study led by graduate student researcher Torrey Rhyne-Carrigg, the group used high-throughput DNA sequencing technology to determine the gene expression activity of approximately 500,000 different versions of the initiator.
ReportedView cited source - [3]
The researchers used machine learning on that data to create an AI model that decoded the initiator's signature DNA base sequence pattern.
ReportedView cited source - [4]
With the decoded pattern, the researchers searched for the initiator's telltale sequence and found that about 60% of human genes contain the initiator.
ReportedView cited source - [5]
Kadonaga said the AI models were found to provide, for the first time, strong predictions of the presence or absence of the initiator in human genes.
- [6]
The announcement describes the models' predictions as 'strong' but reports no accuracy statistic, benchmark comparison or held-out test result.
ReportedView cited source
Sources & coverage · 2 publishers
The reporting this story was synthesized from, earliest first. Every link goes to the original.
- phys.orgyesterdayAI decodes DNA initiator sequence found in about 60% of human genes
- sciencedaily.com8h agoA hidden “on switch” in human DNA has finally been decoded
Additional citations
- James T. Kadonaga, UC San Diego, quoted in the announcement
- James T. Kadonaga, quoted in the announcement


