Science1 publisher2 min readPublished
Salk catalogs 1,067 uncharacterized microproteins in the human frontal cortex
Standard gene models leave out short open reading frames, so a Salk team built a search database that includes them and pointed it at existing proteomics data from postmortem frontal cortex samples.
The Scientist · Science desk

What happened
- Salk Institute researchers published what they describe as the first microprotein atlas of the human frontal cortex with and without Alzheimer's disease, in Nature Aging.
- The paper reports 1,067 previously uncharacterized microproteins with high-confidence spectral support, none of them present in reviewed UniProtKB entries.
- The material came from hundreds of postmortem frontal cortex samples from donors with and without Alzheimer's, including the Religious Orders Study/Memory and Aging Project cohort.
- Deleting the gene that makes that microprotein in microglia impaired mitochondrial respiration, the paper's evidence that it matters for how those immune cells make energy.
Compiled by The ScientistSomething wrong?How this is made
Why it matters
- constraint The gene models drop short open reading frames before any search starts, so a target list assembled from canonical proteome annotation will miss this class of molecule however deep the proteomics goes.
- capability Any group already holding brain mass spectrometry data can re-search it against these sequences, without new tissue and without new instruments.
- decision Anyone mining the list has to judge, locus by locus, whether a detected microprotein is doing biology or merely flagging a gene whose transcription or splicing is disturbed.
- exposure Where the unannotated product is the dominant one, published effects credited to a locus's canonical protein become open to re-examination.
Microproteins come from small open reading frames and typically run to 150 amino acids or fewer, which is part of why they have been hard to detect and study [1]. The obstacle Alan Saghatelian describes sits upstream of the instrument. "Every reference proteome is built on gene models that exclude smORFs by construction," he told GEN. "So the first motivation was straightforward: build a search database that can actually see these sequences, and point it at the deepest human brain proteomics data that exists." [13]
The atlas combines transcriptomics, mass spectrometry and deep-learning-predicted spectra [4], searched with custom tools including ShortStop, an AI-powered microprotein finder developed in Saghatelian's lab [7]. Brendan Miller, the first and co-corresponding author, said the team was "able to take all these technologies and tools and reapply them to existing data from nearly 500 brains to find new microproteins" [8]. The team published the database for other labs to download. "We were able to create an entirely new database that researchers can download and use to better interpret functions of genes," Miller said [9].
MKKS is the locus the paper takes furthest. A small open reading frame there encodes a 63-amino-acid microprotein that appeared to be the predominant translation product at that locus, and it is downregulated in Alzheimer's disease [11]. Knocking out the microprotein-making gene in microglia impaired mitochondrial respiration [12]. That knockout was run in microglia and measured oxygen consumption. The postmortem tissue was sampled once, so whether the drop measured in Alzheimer's brains helps cause the disease or follows from it is untested.
Saghatelian volunteers the ambiguity in his own list. There are two ways to read an expressed microprotein, he said: it may be a marker, evidence that its gene's transcription or splicing is disrupted, without the peptide itself doing anything, or it may be a bioactive molecule with biology distinct from the canonical product at that locus [15]. Of the 1,067 uncharacterized microproteins the paper reports [5], GEN's account describes a functional test for one, about 0.09 percent of them [17].
The disease signal in the atlas runs in both directions. Some microproteins were expressed differently in Alzheimer's samples, and the paper reports that Alzheimer's cells tended to show higher overall microprotein expression [10], while the MKKS product went the other way [11].
For loci where an unannotated product dominates, the consequence falls on experiments already run and published. "The general implication is uncomfortable: the most abundant and most tissue-relevant protein product at a locus can be the one that isn't annotated," Saghatelian said [14].
What to watch
- Whether independent labs re-searching their own brain mass spectrometry data recover the same 1,067 entries.
- Samples staged by pathology, or an animal model, that show when the MKKS microprotein falls relative to symptoms.
- Whether a second entry from the atlas gets a knockout, and whether it also hits mitochondrial respiration.