Unbiased headlines. Facts, not spin.
Every story is an unbiased news briefing written from 113+ sources across the spectrum — sources linked so you can verify it yourself.
Oxford Fed 125,000 Bodleian Scans Into OpenAI's Training Set. Its 2025 Announcement Never Said That Part

A library deal that turned out to be two deals
The University of Oxford announced a partnership with OpenAI in March 2025. The pitch was simple: OpenAI's software would help digitize centuries-old texts from the Bodleian Libraries, making fragile manuscripts searchable and available to students and researchers who'd otherwise need to fly to Oxford and request them in person.
What the announcement didn't say, according to internal documents obtained by The Guardian through a freedom of information request, is that the same scans were being used to "populate the OpenAI training set." Digitizing a book makes it readable online. Feeding it into a training set teaches a commercial AI model to write like it, mimic it, and generate text patterned on it.
By June 2025, Oxford had shared 125,000 scanned images with OpenAI, according to The Guardian and corroborated by Newsbytes and The Next Web. The material includes PhD theses written at European and American universities in the 19th and 20th centuries, plus a rarer haul: 10,000 16th-century "broadside ballads," song lyrics and musical notation once handed out on Tudor street corners. Staff have also discussed digitizing 18th-century Irish state papers, the private letters of novelist Maria Edgeworth, and Dorothy Hodgkin's penicillin research notebooks, per The Guardian.
The pilot, the money, and the classroom rollout
According to Crypto Briefing, the pilot project that kicked off in February 2025 targeted roughly 3,500 public-domain dissertations dating from 1498 to 1884. The deal runs five years and falls under OpenAI's NextGenAI program, which has committed $50 million in research grants, compute, and API access across a small group of elite institutions. Oxford is the only UK member; the others are Boston Public Library, Caltech, MIT, and the University of Michigan, per The Guardian and Inkl.
On the campus side, Oxford began rolling out ChatGPT Edu university-wide starting in September 2025, after piloting the tool with about 750 users, according to Crypto Briefing. Open-access reports on the digitization project's findings are expected by early 2026.
Staff raised the flag. Oxford says it never hid the ball
Meeting minutes from the Bodleian governance committee, also obtained via FOI, record staff worries about reputational risk from tying a 900-year-old institution to a company that has faced years of copyright and data-sourcing criticism under CEO Sam Altman. Staff also questioned how a deal with an energy-intensive technology squares with Oxford's own environmental commitments, according to The Guardian and Archyde.
A public university partnering with a for-profit AI company and not spelling out in its own press release that the material would train a commercial model looks like burying the lede. Readers and students deserve to know what's actually happening with the institution's archives, not a sanitized version of it.
Oxford disputes that framing. A university spokesperson told The Guardian the amount of material digitized is "modest in scale," limited strictly to out-of-copyright works, and shared on a non-exclusive basis, meaning OpenAI doesn't get sole rights. The Bodleian retains rights to the scans and says it will start publishing them openly online within months. The spokesperson also said, per The Next Web, that scanning was Oxford's primary goal but that staff had been open internally that the texts would also serve as training data, meaning the university's position is that this was known, just not headlined.
OpenAI, for its part, framed the arrangement as a public good. "With more than a billion people using this technology in everyday life, it's important it reflects different cultures, histories and perspectives," a company spokesperson told The Guardian, adding the firm was "proud" to help preserve historical knowledge inside its models.
Why the scramble for old books
Scraped internet data is increasingly polluted with AI-generated slop, making it less useful for training new models, according to The Guardian and The Next Web. That's pushing AI firms toward physical archives, and secondhand booksellers have noticed: Superpower Daily and The Guardian both report a spike in orders for obscure titles, like an 18th-century guide to African agricultural implements, on the theory that books never digitized are the last clean data left. In August 2025, 404 Media reported tracking a box of rare books to an Amazon-linked operation that scans and destroys physical copies for AI training. The Bodleian's books stay intact under its deal.
What's still open
The legal question here is narrow: the scanned dissertations and ballads are public domain, so no copyright is being violated. The unresolved question is disclosure. Oxford's contract opens the door to digitizing its full 23-million-item collection, per The Guardian and Newsbytes. Whether future phases of that project get the same public framing as the first one, or whether Oxford spells out the AI-training use up front, is something students, faculty, and UK taxpayers funding the institution have a right to watch for when those open-access reports land in early 2026.
Sources used for this briefing
This briefing was written by UBH's AI agent — these are the reporting inputs it draws on, linked so you can verify.