Skip to content

Product1 publisher3 min readPublished

Austria's Ancient Greek model proposes missing words for scholars to choose between

The Austrian Academy of Science built Apollo with Mistral and Sail Reply on roughly 600 million words of historical Greek, and made it free to academics through a chatbot. A scholar still picks the word that goes into the record.

The Product Desk · Product desk

Illustration accompanying Austria's Ancient Greek model proposes missing words for scholars to choose between

What happened

  • The Austrian Academy of Science releases Apollo on Wednesday, billed as the world's first advanced large language model for Ancient Greek and built in partnership with Mistral and Sail Reply.
  • The model was trained on roughly 600 million words of historical Greek drawn from manuscripts, papyri and inscriptions.
  • Academics can use it free of charge through a chatbot interface.
  • The papyri still awaiting restoration are largely mundane documents: personal letters, marital contracts and civil service papers.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

  • constraint Because the tool stops short of a finished reading, the scholar's attention stays the limit on how many fragments get restored. A department cannot plan headcount around it.
  • contradiction Apollo is presented as a national academy's model, and a commercial AI lab co-built it. Another archive holder can copy the curated corpus and the review step.
  • precedent Sail Reply's Vlitas says the method could move to Latin, Egyptian or any field with a large corpus, so institutions sitting on unindexed archives can expect the same proposal.
  • exposure Whoever accepts a suggested word owns it in the published edition, and later scholarship cites the edition rather than the chatbot that produced the shortlist.

Ancient Greek was written without spaces between words. So before anyone can guess at a missing word in a torn papyrus, the word divisions have to be found, the document dated, its socio-political context weighed, and reference works consulted [6]. Four separate acts of specialist judgement come before the guess itself [7]. Stephen Colvin is a professor of classics and historical linguistics at University College London. He said of the people who can do all of it: "There are very few people in the world who are that good at Greek history" [8].

Anna Dolganov, a historian and papyrologist at the Austrian Academy of Science, said the model matches the register of whatever it is shown: "When it sees Homer, it supplements Homeric Greek. When it sees an inscription in Doric dialect, it uses Doric dialect" [9]. The WIRED report does not include an accuracy rate for those suggestions or a comparison against a general-purpose model [18]. The case for a narrow model here therefore rests on the corpus that went in and on the shape of what comes back out.

What comes back out is deliberately unfinished. Apollo offers a selection of candidate words and the scholar selects one, a choice the project made to keep probabilistic guesses from polluting the historical record with errors [10]. "The crucial point is that human competence needs to remain," Dolganov said [11]. Armand D'Angour, a professor of classical languages and literature at the University of Oxford, put the same design from the user's side. "If I had a machine telling me, 'Here are the three possible words that could fit into that gap,' it would speed up matters considerably," he said [13].

The user is a specialist working through ordinary paperwork. Academic libraries hold hundreds of thousands of Ancient Greek papyrus fragments, many of them so damaged that their meaning is probably lost [5]. The ones still waiting are personal letters, marital contracts and civil service papers [15]. Colvin said: "If you were a layperson, you might think suddenly we'll get a few new plays by Sophocles, but that's not going to happen" [14]. D'Angour said: "Every time something is produced, it adds a tiny element of knowledge about the ancient world" [20].

Two things decide whether this shape transfers to a domain model in your own field. Do you hold a corpus nobody else has indexed, and does a named person accept or reject each suggestion before anyone else sees it? Dimitris Vlitas, a partner at Sail Reply, said the same technique could be applied to Latin or Egyptian, or any academic discipline that would benefit from the distillation and indexing of a large corpus [17]. Those disciplines come with editors and peer reviewers already in the loop. A team whose model output goes to a customer with nobody in between cannot copy the shortlist design. There is no one there to choose.

What to watch

  • Whether the Academy publishes accuracy rates for Apollo's suggestions, or any comparison with a general-purpose model, once academics are using the chatbot.
  • Whether Vlitas's Latin or Egyptian version appears, and which institution owns the corpus in that case.
  • Whether classics journals begin asking authors to disclose when a reconstruction came from a model suggestion.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories