Oxford University allowed OpenAI to add Bodleian Library scans to its AI training data, according to internal papers reported by the Guardian on Saturday by Ethan Penny and Dan Milmo. By June 2025 the Bodleian had sent OpenAI 125,000 scans of old PhD theses from European and US universities. The March 2025 public announcement had described only digitization for students and scholars, with no mention of training. The disclosure matters for business because licensed library collections are becoming a foundation for commercial models.
How the Bodleian scanning deal worked
The partnership was disclosed in March 2025 as a project to scan rare texts and open them to more students and scholars. Internal papers seen by the Guardian indicate that material scanned by OpenAI at the Bodleian entered training data. By June 2025 the library had transferred 125,000 scans of old doctoral theses dating from the 19th and 20th centuries and originating at European and US universities. Oxford is the only UK member of OpenAI's NextGenAI group, which also includes Boston Public Library, Caltech, MIT and the University of Michigan.
Oxford describes the transfer as small in scale, out of copyright and non-exclusive, with the library retaining rights to the scans. The institution plans to begin publishing the scans online in the next few months. Under the arrangement the physical volumes remain intact, with scanning done at the library rather than through removal or destruction. An OpenAI spokesperson said the training use had not been concealed, noting that digitization was the main goal while staff had been open about model training.
The arrangement fits demand for printed, human-written material as the web fills with AI-generated text. Some buyers of books for training cut volumes apart for scanning, which has upset secondhand booksellers. In August, 404 Media traced a box of rare books to an Amazon scanning site that destroys books for AI. Notes from Bodleian staff meetings, obtained through a freedom of information request, show concerns about reputational risk and AI energy consumption. OpenAI told the Guardian that with more than a billion people using the technology, systems should reflect different cultures, histories and perspectives.
What this means for companies using AI
For companies that buy or build on commercial models, licensed archival text can improve factual grounding and breadth of reference compared with web-only data. Customer support agents, research assistants and internal search tools gain from older academic writing that is structured, cited and free of recent AI-generated noise. A small firm feels the effect as better answers without negotiating its own library deals. A large organization with compliance duties gains a clearer provenance trail, since the material is described as out of copyright and non-exclusive.
The case also shows what licensed data does not guarantee by itself. The collection covers 19th- and 20th-century theses, not current commercial, technical or customer documentation, so recency and domain fit still need checking. Disclosure timing matters: the March 2025 statement emphasized access, while training use emerged later in internal papers, which underlines the need to ask vendors directly. Buyers should ask which collections trained a specific model version, what rights and exclusivity apply, and how updates handle corrections.
A concrete marker to watch is the Bodleian's planned online release of the scans in the next few months. The scope, quality and usage terms of that publication will show whether the project delivers public research value alongside training benefit. Further additions to NextGenAI will indicate how fast licensed library sourcing scales for commercial AI.
