DocsQuick StartAI News
AI NewsOpenAI Puts Oxford’s Ancient Books into Its Training Dataset
Industry News

OpenAI Puts Oxford’s Ancient Books into Its Training Dataset

2026-09-26T18:07:21.617Z
OpenAI Puts Oxford’s Ancient Books into Its Training Dataset

The Bodleian Libraries partnership originally advanced by the University of Oxford under the banner of digitization has been confirmed by internal documents to involve an OpenAI training dataset. The controversy centers not on whether ancient books can be used to train AI, but on the boundaries of authorization, disclosure of information, and the distribution of returns from public cultural assets.

Digital Collaboration Ultimately Became Training Data

On September 26, it was disclosed by British media and subsequently reported by domestic media that the University of Oxford had allowed OpenAI to use part of the Bodleian Libraries’ historical holdings to build a model training dataset.

The key point is this: when the University of Oxford announced the collaboration in March 2025, it emphasized the digitization of documents and making the collections more accessible to students and researchers. However, internal documents revealed today explicitly state that content digitized by OpenAI had been used to “build OpenAI’s training dataset.”

This is not a minor difference in wording. For a library, “asking a technology company to help scan documents” and “allowing a technology company to put those documents into a training dataset” correspond to two entirely different systems of authorization, governance, and accountability.

The Bodleian Libraries are the University of Oxford’s principal research libraries and the second-largest library system in the United Kingdom after the British Library. Their holdings include not only ordinary publications, but also manuscripts, dissertations, private correspondence, laboratory notebooks, and rare historical documents. These materials may not rival the internet in sheer volume, but they are scarce, reliable, and structurally complete—precisely the kind of data model developers want most at this stage.

A collage of the exterior of the University of Oxford’s Bodleian Libraries and pages from historical-document scanning

As of June 2025, the Bodleian Libraries had provided OpenAI with approximately 125,000 scanned facsimile pages related to historical dissertations, covering doctoral theses from universities in Europe and the United States from the 19th and 20th centuries. The materials already scanned also include approximately 10,000 16th-century “broadside ballads,” preserving lyrics and musical notation that circulated on the streets during the Tudor period.

The additional digitization projects discussed by the two parties include:

  • The Irish national archives from the 18th century;
  • The private correspondence of Irish novelist Maria Edgeworth;
  • The laboratory notebooks of Dorothy Hodgkin, a Nobel Prize winner in Chemistry, concerning her research on penicillin.

If these materials enter the training pipeline, their significance clearly goes beyond “a few more books.” They could fill major gaps in the historical language, archival writing, early scientific records, and local cultural expression that are extremely sparse in mainstream internet corpora.

What Is Truly Scarce Is Not Data Volume, but “Clean Data”

Over the past several years, large models have rapidly expanded their training scale by relying on web pages, code repositories, e-books, and forum content. However, high-quality natural-language data from the public internet is becoming increasingly difficult to obtain.

On the one hand, high-quality publishing organizations are introducing paywalls, restricting crawlers, or signing licensing agreements directly with model companies. On the other hand, search results and social platforms are being heavily polluted by AI-generated content. If models continue indiscriminately scraping the web, they can easily ingest material generated, rewritten, or even fabricated by other models.

This is creating an increasingly realistic problem: the total volume of training data continues to grow, but corpora with verifiable sources, clear dates and authorship, and careful human cataloging and organization are becoming increasingly expensive.

This is where the value of libraries and university archives becomes apparent. They generally provide data that is not made up of web fragments scraped on the fly, but that has the following characteristics:

  1. Clear provenance. The author, date, edition, and collection location of a document can usually be traced;
  2. Scarce content. Many manuscripts, papers, and local historical materials have not been adequately incorporated into the public internet;
  3. Lower noise. Content cataloged by libraries and filtered through academic processes generally has a higher average information density;
  4. A broad time span. Historical documents can reduce the tendency of training datasets to become overly concentrated in the contemporary English-language internet;
  5. A relatively assessable rights status. Some ancient works have already entered the public domain, but scans, databases, and private archives may still be subject to contractual restrictions, database rights, or privacy rules.

For OpenAI, this kind of collaboration is more like building an “institutional data supply chain” beyond public internet data. It may not significantly increase the total number of tokens, but it could improve the cultural coverage and credibility of the training set.

However, outsiders currently do not know which specific stage of training these holdings entered, or whether they have already been used in any officially released model. “Building a training dataset” could encompass pretraining, continued pretraining, OCR correction, document understanding, retrieval evaluation, or other data-processing stages. Based on the existing documents alone, it is impossible to directly assert that any particular GPT model has absorbed all of these materials.

This is a technical boundary that must be preserved in reporting.

Scanning Ancient Books Does Not Mean Models Can Directly Understand Them

Digitizing historical documents does not end with photographing the pages.

For models to actually use these materials, they typically must go through image cleanup, layout analysis, text recognition, human proofreading, metadata association, deduplication, and quality screening. Sixteenth-century ballads, handwritten letters, and early laboratory notebooks are particularly difficult: their typefaces are inconsistent, their spelling differs significantly from modern language, and their pages may contain annotations, charts, stains, and binding shadows.

Traditional OCR is good at processing modern printed material with regular layouts, but its error rate rises rapidly when it encounters historical typefaces and complex pages. Multimodal models can process page images directly, but they still need high-quality samples to understand text regions, reading order, and the relationship between text and images.

Therefore, what OpenAI may obtain from projects of this kind is not only the documentary content, but also a set of highly valuable data-engineering outputs:

  • Original high-resolution scanned images;
  • OCR-transcribed text and correction results;
  • Collection metadata such as authors, dates, and subjects;
  • Annotations of page layouts and text locations;
  • Correspondences among images, text, and catalog records.

From the perspective of model capabilities, these data can help improve historical-document recognition, archival search, and academic question answering. But their immediate impact on general-purpose models should not be overestimated. Compared with training corpora containing trillions of tokens, hundreds of thousands of document pages is not a particularly large volume. The data are more likely to improve certain long-tail fields than to produce a sudden overall leap in model capabilities.

In other words, the strategic value of this dataset is greater than its benchmark value.

The Biggest Controversy Is Not Copyright, but the Boundaries of Notice and Authorization

Because ancient books are so old, much of their content may already be in the public domain. Some people therefore believe that since anyone can read them, there should be no problem using them to train models.

It is not that simple.

The public domain primarily resolves the copyright status of the original works. It does not automatically answer the following questions:

  • Can high-resolution scans produced by a library be used in bulk for training by a commercial organization?
  • Do private letters and modern academic materials contain personal or sensitive information?
  • Did donors’ original authorizations cover machine-learning uses?
  • Can OpenAI retain, reprocess, or combine the materials with other data over the long term?
  • Do the resulting models and commercial products need to provide anything in return to public institutions?
  • Can researchers obtain access to the data on the same terms as OpenAI?

The issue of disclosure is even more sensitive. The University of Oxford initially emphasized “digitization” and “convenient access.” Ordinary members of the public could easily understand this to mean scanning, preservation, and the creation of a search system—not training a commercial large language model.

If the contract permitted training from the outset but the public announcement deliberately downplayed that use, the issue is one of transparency. If training was added as the project progressed, then further questions must be asked: Who approved the change in use? Were the ethical and legal assessments repeated?

The University of Oxford has denied concealing the machine-learning use from the public and students. However, internal meeting minutes show that faculty and staff, including members of the Bodleian Libraries’ governing committee, had already expressed concerns about reputational risk. This indicates that the controversy is not coming solely from outside public opinion; there is also no unreserved consensus within the university.

“Preserving Culture for Future Generations” Is Only Half the Truth

OpenAI’s response was that the company was honored to help “today’s AI models safeguard humanity’s historical and cultural treasures for future generations.” It also emphasized that more than one billion people around the world already use related technologies in their daily lives, and that models therefore need to reflect different cultures, histories, and perspectives.

This argument is not without merit.

Large quantities of historical material have long been locked away in physical storage facilities, requiring ordinary researchers to apply for access in person. Having companies with greater funding and technical capabilities bear the costs of scanning, recognition, and organization could indeed accelerate the digitization of cultural heritage. If the final results are made available to academia and the public, the collaboration could have clear public value.

But “preserving culture” and “training commercial models” cannot automatically be treated as equivalent.

Whether the digitized results are made available, the level of detail at which they are released, what reuse rights the model company possesses, and whether the library can reclaim the structured data—these are the factors that determine whether the collaboration is truly a public cultural project or a data acquisition deal conducted in exchange for digitization.

At a minimum, a more reasonable collaboration should disclose several categories of information:

  • The scope of the documents scanned and used for training;
  • The copyright, privacy, and opt-out mechanisms governing the data;
  • How long OpenAI will retain the scans and derived data;
  • How the digitized results will be made available to the public and researchers;
  • Whether third-party audits of data usage will be permitted;
  • What return public institutions will receive if the resulting models or products generate commercial revenue.

If these conditions are not transparent, “AI helping to preserve history” can easily become an attractive layer of packaging covering commercial uses.

The High-Energy-Consumption Controversy Cannot Be Dismissed in a Single Sentence Either

According to the disclosed meeting minutes, Oxford faculty and staff were also concerned that the energy-intensive technology involved in the OpenAI collaboration could conflict with the university’s own environmental commitments.

It is important to distinguish between the energy consumption of scanning and OCR and that of training a large-scale model; they are not on the same order of magnitude. The scanning of a particular group of documents cannot simply be equated with the training of an entire model, and it is difficult to accurately calculate how much carbon emissions were caused by this dataset alone.

Nevertheless, the university’s concern is reasonable. As universities set net-zero emissions targets while providing training resources to model developers with enormous computing demands, they do need to explain how the environmental costs are being assessed. Especially when a collaboration is described as serving the public interest, energy use, server locations, and emissions-reduction measures should not be completely hidden behind commercial confidentiality clauses.

OpenAI Is Turning Universities into New Data Gateways

The Bodleian Libraries are not an isolated case. Under the NextGenAI project, OpenAI has already entered into similar collaborations with the Boston Public Library, the California Institute of Technology, the Massachusetts Institute of Technology, and the University of Michigan, among other institutions. The University of Oxford is the project’s only British member.

This reveals a new stage in competition among large models: in previous years, the contest was over who could scrape more public web pages; now it is over who can obtain closed collections, specialized databases, and institutional archives that others cannot replicate.

Compared with continuing to expand crawler operations, signing direct agreements with universities and libraries offers three advantages: a clearer chain of rights, more stable data quality, and greater difficulty for competitors seeking access to the same materials. This is essentially part of the same trend as model companies purchasing licenses for news content and connecting to publishers’ book collections—high-quality corpora are shifting from public resources presumed to be freely scrapeable into strategic assets that require negotiation.

This could also widen the advantage of leading model companies. Organizing and digitizing university archives is expensive. Only companies with sufficient funding and the ability to provide scanning technology and long-term collaboration plans can easily gain access to these institutions. Even if smaller model teams know where the data is, they may not be able to obtain it.

The Verdict: The Collaboration Has Value, but Default Opacity Is No Longer Viable

From the perspective of technology and cultural preservation, OpenAI’s participation in digitizing the Bodleian collections is not inherently a bad thing. Large quantities of historical documents do need better scanning, recognition, and retrieval tools, and models also need to reduce their overreliance on the contemporary English-language web.

The problem is that public and academic institutions cannot use “digitization” as a blanket term for all subsequent uses. Scanning, online display, academic search, training OCR models, and training general-purpose commercial large models are different types of processing activities. They should be explained separately rather than placed under a broad authorization.

For developers, this episode also sends a more direct signal: future differences in model capabilities may not come solely from architecture and computing power, but also from the exclusivity of training data. As foundation models become increasingly similar, whoever can obtain high-quality, traceable, professionally specialized data that has not been contaminated by the internet is more likely to build barriers in vertical tasks such as historical research, law, medicine, and scientific research.

What OpenAI has obtained is not a few boxes of old books, but a channel into institutional knowledge repositories. The questions truly worth asking are not whether ancient books can be fed to AI, but who has the authority to decide, how much the public is told in advance, and who ultimately receives the value created from public cultural assets.

Sources

Related Articles

View All

Contact Us

We usually reply quickly during business hours

Scan WeChat

Support: Hub Assistant

WeChat ID: