Skip to content

Leadership1 publisher2 min readPublished

Oxford's Bodleian scans fed OpenAI's training set, internal documents show

Oxford let OpenAI train on Bodleian texts it digitised, with 125,000 dissertation images shared by June 2025, internal documents show. Oxford says staff never hid the training use, though its 2025 announcement spoke only of wider access.

The Board Room · Leadership desk

Photograph accompanying Oxford's Bodleian scans fed OpenAI's training set, internal documents show
Photo: ox.ac.uk

What happened

  • By June 2025, 125,000 images scanned from 19th- and 20th-century dissertations had been shared with OpenAI from the Bodleian collection.
  • Minutes released under freedom of information show staff, including Bodleian governance committee members, worried about reputational risk and the university's environmental commitments.
  • Oxford says the material is out of copyright, that the Bodleian keeps the rights to the scans, and that it will begin publishing them openly within months.

Compiled by The Board RoomSomething wrong?How this is made

Why it matters

  • contradiction Oxford says staff were open about the training use while its public announcement left it out; both can hold. The public announcement omitted what staff discussed internally, and that omission is the reputational risk its own minutes flagged.
  • exposure Terms agreed for a first tranche Oxford calls modest could extend to a 23m-item collection and an "Ask the Bod" chatbot. The disclosure choices made now would carry over to that much larger transfer.
  • decision As scraped web text loses value for training, collections held only in print gain bargaining weight, and library boards must decide whether training rights are part of a scanning deal or a separate term.
  • precedent As the only UK member of NextGenAI, Oxford is likely to be the reference case when OpenAI or its rivals approach other British libraries.

Oxford's own account of the deal is the version a board would sign off. "The material digitised through the project with OpenAI is modest in scale, out of copyright, and OpenAI's use of the material is not exclusive," a university spokesperson said [12]. The Bodleian keeps the rights to the scans and will begin publishing them openly online within months, the spokesperson said [10]. Those are real protections, and they cover what Oxford retains. The internal documents record what OpenAI gets: the digitised material is being used to "populate the OpenAI training set" [2].

A library holding texts that exist nowhere online now holds something model developers want [16]. Scraped websites are increasingly saturated with AI-generated material, making them less useful for training, and developers have turned to physical, often historical, book collections, according to the Guardian [16]. Anthropic has spent tens of millions of dollars buying books, slicing off their spines to scan them and then pulping them, though it says it does not buy and destroy rare and antiquarian books [17]. Under Oxford's deal the Bodleian's collections stay intact [19].

Meeting minutes obtained under freedom of information show staff, including members of the Bodleian governance committee, raising concerns about reputational risk and about the deal's effect on the university's environmental commitments [5]. The March 2025 announcement said the project would make the content more widely available to students and researchers [3]. It did not say the material would be used to train OpenAI's models [4]. Oxford rejects the suggestion that this was hidden. Its spokesperson said digitisation was the university's primary interest and that staff had been open that the project would also contribute training data [11].

"Modest in scale" describes the first tranche. By June 2025, 125,000 images scanned from 19th- and 20th-century dissertations had been shared with OpenAI [6]. Other scans include a collection of 10,000 16th-century broadside ballads [7]. Staff have discussed digitising 18th-century Irish state papers and Dorothy Hodgkin's penicillin notebooks [8]. The minutes also discuss an "Ask the Bod" chatbot, and the contract raises the prospect of mass digitisation of the Bodleian's 23m items [9].

Oxford is the only UK member of NextGenAI, the OpenAI project under which the company has struck similar agreements with Boston Public Library, Caltech, MIT and the University of Michigan [15]. OpenAI presents the work as preservation. A spokesperson said the company was "proud" to ensure "the AI models of today preserve the world's historical knowledge for the future" [13]. "With more than a billion people using this technology in everyday life, it's important it reflects different cultures, histories and perspectives," the spokesperson added [14].

What to watch

  • Whether Oxford or OpenAI publishes the contract's terms on training use and on any wider digitisation of the Bodleian's 23m items.
  • Whether the Bodleian begins publishing the scans openly online within the months its spokesperson promised.
  • Whether Boston Public Library, Caltech, MIT or the University of Michigan disclose that material from their NextGenAI agreements also went into OpenAI's training set.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories