Original briefings. Zero spin.
Every story is an original briefing written from 110+ sources across the spectrum — sources linked so you can verify it yourself.
Courts Say AI Companies Can Train on Books. They Still Can't Steal Them.

Training an AI model on a copyrighted book, without the author's permission, is legal in the United States. That's the actual rule right now, set by a federal judge. But it comes with a massive asterisk, and that asterisk just cost one company $1.5 billion.
Judge William Alsup ruled in the case Bartz v. Anthropic that training the Claude chatbot on published books was "exceedingly transformative" and protected as fair use under Section 107 of the Copyright Act, according to Forbes. He compared what an AI model does to a writer studying earlier authors: absorbing style and structure without copying the work itself.
AI companies won their case in part, but not entirely.
Alsup split the case into two separate questions, according to Forbes and Startup Fortune. Training on lawfully obtained books is permitted. Building a permanent library out of pirated copies downloaded from shadow sites like Library Genesis and Pirate Library Mirror is not, and is not protected by fair use.
On July 20, 2026, Judge Araceli Martínez-Olguín granted final approval to Anthropic's $1.5 billion settlement with the authors and publishers who sued, according to Startup Fortune, citing TechCrunch's reporting. That works out to roughly $3,000 per covered work across an estimated 500,000 works, per the Authors Guild figures cited by Forbes. Anthropic also has to destroy the pirated copies covered by the deal.
For a company Reuters reports is projecting $190 billion to $200 billion in 2028 revenue, per Startup Fortune, $1.5 billion is a manageable cost of doing business. Attorney Cathy Gellis told TechCrunch that's exactly why she sees the ruling as a net win for AI companies generally: "Copyright law hinges on copying, but it doesn't hinge on using the work or experiencing the work, consuming the work, reading the work."
Jason Henderson, an IP attorney at JWL International, told TechCrunch the courts are still catching up. "Everybody is very worried right now because the law is all over the place," he said, noting copyright law hasn't been rewritten since 1976, decades before anyone imagined a chatbot ingesting a library's worth of text.
The fair-use question isn't settled across the board. In Thomson Reuters v. Ross Intelligence, a different judge, Stephanos Bibas, ruled the opposite way, finding that Ross Intelligence's use of Thomson Reuters' content to build a competing legal-research AI was not fair use, according to CryptoRank. The distinguishing factor courts keep circling back to is whether the AI product directly competes with the original work's market. Train a general chatbot on novels, and courts have leaned toward fair use. Train a legal-research tool on another company's legal database to sell a rival legal-research tool, and courts have leaned the other way.
That distinction is now the live legal question in a separate case. Publishers Hachette, Cengage, and Elsevier, along with novelist Scott Turow and the group S.C.R.I.B.E., filed a class action against Google on July 10, 2026, in the Southern District of New York, accusing the company of training Gemini on copyrighted books and journal articles without rights to do so, according to Startup Fortune.
Now a separate fight: are AI companies destroying books on purpose?
More than a dozen consumer and public-interest groups, including the Demand Progress Education Fund, the Consumer Federation of America, and the Institute for Local Self-Reliance, sent a letter Friday to FTC Chairman Andrew Ferguson and Commissioner Mark Meador, according to CBS News. Their claim is that AI companies are buying physical books in bulk, scanning them for training data, and then destroying the copies, sometimes destroying what the groups say are among the last surviving copies of certain works.
The groups argue this could violate Section 5 of the FTC Act, which bars unfair methods of competition, according to CBS News. Their reasoning is that if only the biggest, richest AI labs can afford to buy and scan massive book collections, and those originals then vanish, smaller competitors and the public lose access to source material the incumbents already digitized. No FTC investigation has been opened as of this reporting. This is an allegation in a letter, not a finding.
CBS News also reported that 404 Media, a digital publisher, found this week that Amazon is buying books in bulk, scanning them to train its AI tools, and then destroying them. Amazon did not respond to a request for comment on that report, according to CBS News. Anthropic likewise did not respond to CBS News's request for comment on the FTC letter.
The AI industry would likely argue that scanning and discarding a physical book isn't necessarily "destroying" the work in any copyright sense, since the content itself survives in digitized form. Companies could argue bulk-purchasing books they intend to scan is a legitimate way to source training data legally, rather than pirating it, which is exactly what the Anthropic case punished them for not doing. Whether that practice also amounts to anticompetitive hoarding, as the advocacy groups claim, is a separate and unresolved question the FTC has not addressed.
What's left unresol
Sources used for this briefing
This briefing was written by UBH's AI agent — these are the reporting inputs it draws on, linked so you can verify.