Science1 distinct publisher3 min readUpdated
Siddhartha Jayanti derived technical terms from Sanskrit to write the first modern computer science paper in an Indian language. The method is built to be reused.
The Scientist · Science desk

Compiled by The ScientistSomething wrong?How this is made
Siddhartha Jayanti, now an assistant professor of computer science at Dartmouth, has posted a paper on arXiv describing how he manufactured the Telugu technical vocabulary he needed before he could write up his own results [7]. The order of operations is the point: the words did not exist, so he derived them, and the derivation method is reusable across languages.
The underlying research is not linguistics. Jayanti analyzed how groups of processors coordinate by tracking what each one knows and how they communicate, and established fundamental mathematical limits on how quickly multiple cores can safely coordinate and share common data [2]. When he sat down to write it in Telugu, a language with roughly 100 million speakers [1], there was no settled term for something as basic to the work as "shared-memory multiprocessor" [3]. That is the barrier in concrete form: not a shortage of speakers, and not a shortage of results, but a missing lexicon standing between the two.
His fix was construction rather than transliteration. For "asynchronous" - which in distributed computing means processors work independently without waiting for other systems to stay in sync [13] - he built the Sanskrit asamakalika from a- (not), sama (same) and kala (time), then adapted it to asamakalikamu to fit Telugu grammar [12]. In the paper he notes that English got there the same way, assembling "asynchronous" from the Greek a-, syn- and khronos [14]. "Shared-memory multiprocessor" became samvibhakta-smrti bahusamsadhakamu, his most complex example, produced by modifying Sanskrit roots with prefixes and suffixes [15].
The reason to route through Sanskrit is leverage of a specific kind. Indian languages have long borrowed from classical sources with minor changes to word endings or sound substitutions, which Jayanti says means "you can create a common vocabulary across a slew of Indian languages all at once" [11]. That matters in a country of more than 1.47 billion people with dozens of languages, where Telugu is one of 22 official ones [8][9] and accounts for roughly 7 percent of the population by itself [20]. Jayanti also argues that natively derived terms teach: speakers can relate them to other words and understand them linguistically and scientifically [10]. A borrowed English string carries no such structure.
The second obstacle was mechanical. Keyboards were missing characters, fonts rendered Telugu incorrectly, and typesetting software could not handle the script [16], so he assembled a workflow he calls TeluguTeX to set Telugu text and mathematical notation correctly [17].
The work was not a side project that cost him standing. The Telugu chapter of his dissertation won the 2023 ACM Principles of Distributed Computing Doctoral Dissertation Award [4], the chapter was posted on arXiv in 2022 as the first modern computer science research paper in an Indian language [6], and its technical results have fed follow-on work in distributed and parallel computing, including efficient algorithms developed with collaborators at Princeton and MIT [19]. One detail is worth holding onto: the main results had to be translated into English for reviewers [5].
What to watch is whether the tooling and the terminology get adopted by anyone other than their author. Jayanti's own stated dependency is better fonts, keyboards and typesetting tools [18], which is a funding and standards problem, not a research one. The terminology test is whether Sanskrit-derived coinages spread to other Indian-language communities as he predicts, or stay a one-language artifact. And the reviewer question stands unanswered: a vocabulary is not much use if publication still requires an English version.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
Siddhartha Jayanti wrote the first computer science paper in Telugu, an Indian language with approximately 100 million speakers.
The main results were translated into English for reviewers.
The chapter was published in 2022 on arXiv, an open-access archive of research papers, becoming the first research paper in modern computer science written in an Indian language.
Jayanti says: "When you can derive terms natively, they mean things to people. They can relate them to other words, and they can understand them both linguistically and scientifically."
Jayanti says it has been an enduring practice across Indian languages to take words from classical languages such as Sanskrit and make them their own through minor alterations of word endings or sound substitutions: "This means that you can create a common vocabulary across a slew of Indian languages all at once."
Jayanti found that many software programs did not fully support Telugu, with problems such as missing keyboard characters, fonts that displayed Telugu incorrectly and typesetting software that could not properly handle the script.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Specific but single-sourced and largely self-reported
Every fact comes from one publisher's account, and the load-bearing details (coinages, TeluguTeX, downstream Google usage) are attributed to the researcher himself. The story does contain checkable anchors — a 2022 arXiv publication, a named 2023 ACM PODC award, named collaborating institutions — but no arXiv identifiers, paper titles, links or third-party confirmation are supplied, and the 'first in an Indian language' priority claim is asserted rather than demonstrated.
Proof of concept by one author, with one unverified downstream user
Adoption of the method itself is n=1: a single author, a single Telugu paper, and a personal toolchain with no described distribution. The strongest adoption signal is indirect — algorithms derived from the underlying theory reportedly shipping in Google's open-source graph software — but that is a single unquantified disclosure attributed to the researcher. No other researchers, journals, publishers, tool maintainers or language bodies are shown adopting the vocabulary or the workflow.
Reusability framed ahead of demonstrated reuse
The technical and linguistic work is real and specific, but the framing runs ahead of it in two places: the method is presented as generalizable to a slew of Indian languages when only one language and one paper exist, and the tooling and Google-usage claims are stated without artifacts or numbers. The gap is moderate rather than severe because the concrete deliverables — a published paper, a named award, documented coinages — are not exaggerated.
Researcher-promoted work in a science-PR channel
The account is a science-communication piece built around one researcher's own framing, and that researcher has explicit stakes in the narrative: establishing a 'first', promoting a proposed Samskrtam Technical Lexicon Project, and positioning the approach relative to national education-policy textbook work. Uncomfortable questions a skeptical outsider would ask — acceptance of the coinages by Telugu-language scholars, competing translation approaches, verification of the Google claim — are absent, which is consistent with promotional incentives rather than fabrication.
Moderate on the record, low on the forward claims
Confidence is fair for the checkable historical record (a 2022 arXiv Telugu chapter, the 2023 ACM PODC award, a current methodology paper, named collaborations) because these are specific and falsifiable. It is low for the claims that matter most to a reader deciding whether anything changes: cross-language reusability, tool improvement effects, and third-party production usage. One publisher, no independent corroboration, and no identifiers cap the overall level.
build
Prompt caching cuts agent API costs 41-80%, but only if tool results stay out of the cache1 distinct publisher
build
Grok 4.6 lands in Copilot two days after launch, and the model picker becomes a procurement problem1 distinct publisher
security
Google's reference agent approved a $10,000 refund on a $149 order, on purpose1 distinct publisher
product
Washington's secret AI test is coming for open weights, and release dates go with it2 distinct publishers
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 17, 2026