Science1 publisher3 min readPublished
Patent-likeness scoring leaves the lab, and your abstract becomes the interface
A Sydney firm's Translation Readiness Index ranks papers by how closely their titles and abstracts resemble patent-linked work. It is now being tested with several universities.
The Scientist · Science desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- A machine-learning tool called the Translation Readiness Index (TRI) scores how 'patent-like' a scientific paper is, months or years before any deal, patent filing or spin-off reveals commercial potential.
- TRI performs a linguistic analysis of a paper's title and abstract and measures how similar the paper's vocabulary is to publications that have previously been paired with patents.
- The method was developed by researchers at the data-analytics firm League of Scholars in Sydney, Australia, and posted as a preprint on arXiv.
- The preprint describing TRI has not yet been peer reviewed.
- The team is now testing TRI with several universities.
Compiled by The ScientistSomething wrong?How this is made
Why it matters
Researchers at League of Scholars, a data-analytics firm in Sydney, have posted a preprint describing the Translation Readiness Index (TRI), a model that scores how closely a paper's title and abstract resemble the vocabulary of publications previously paired with patents [1][2][3]. The work has not been peer reviewed [4], but the team says it is now testing the index with several universities [5], which shifts the question from whether the signal exists to who screens with it.
The mechanism is narrow and worth stating plainly. TRI reads titles and abstracts only, and does not directly assess a paper's underlying data or results [6]. It was trained on 20,610 papers, of which 9,431 had been matched to patents [7]; the remaining 11,179 had not [8], making patent-linked work about 46% of the set [9]. Titles and abstracts went into five classifiers, and the best gave a 78% chance of ranking a patent-linked paper above an otherwise comparable paper with no patent link [10]. Put the other way, roughly one such pairwise comparison in five comes out backwards [11]. Papers eventually cited in patents used words such as prototype, device and design more often than papers that were not [12].
To check the top of the ranking against other markers of commercial activity, the team took the 100 highest-scoring papers by University of Western Australia authors published between 2019 and 2026 [13]. According to co-author Paul McCarthy, 83 had industry-affiliated co-authors and 34 involved a UWA-affiliated author who had patented before, both more common than in a random sample [14][15][16]. That is agreement with proxies rather than with outcomes, and the proxies are close cousins of the input: industry collaboration is one of the settings that produces prototype-and-device prose in the first place [12][14].
McCarthy, a computational social scientist and co-founder of the firm, describes the tool as "a new way of triaging or ranking" research and as an estimate of the probability that a paper uses "patent-like language" [17][18]. He does not recommend basing investment decisions on the score alone, because it is a probabilistic ranking, though he says it might uncover unexpected gems [19]. Ben Miles, co-founder of the London early-stage deep-tech firm Empirical Ventures, told Nature the tool could work as an external signal for academics and funders deciding which ideas deserve further support before they are mature enough for investors [20]. Patentable technologies are not always commercially viable, so any such measure will be imperfect for investors' purposes [21].
Adoption is being driven by arithmetic rather than by insight. Haystack, a comparable tool built for Cornell University's technology team, scans up to 13,000 papers a year, more than the technology-transfer staff can inspect manually, according to Matt Marx, who built it and is the university's vice-provost for entrepreneurship, innovation and external engagement [22][23]. TRI is one of several scouting tools offering the same capability [24].
The consequence for researchers is a pricing change in word choice. Because the model never sees the results [6], an author who writes device instead of system, or prototype instead of demonstration, moves the score without moving the science. Watch whether the university pilots report how many high scorers later yield filings, licences or spin-offs [5], whether the preprint survives review [4], and whether any vendor extends scoring past the abstract to full text [6]. Nature notes that League of Scholars has previously worked with the Nature Index as a data provider, and that its article was produced independently of that work [25].