READ. SCROLL. LISTEN.

Original briefings. Zero spin.

Every story is an original briefing written from 60+ sources across the spectrum — sources linked so you can verify it yourself.

← Back to headlines

The Atlantic's AI Music Database Is Now Public and Searchable, Naming Google and Stability AI as Users

The Atlantic's AI Music Database Is Now Public and Searchable, Naming Google and Stability AI as Users
Reporter Alex Reisner has made four datasets totaling more than 21 million songs fully searchable to the public. Google and Stability AI have both confirmed using at least some of the data in published research papers. The core legal question, whether scraping licensed music from YouTube and Spotify for commercial AI training violates copyright law, remains unanswered in court.

Atlantic reporter Alex Reisner has built a publicly searchable database so anyone can look up whether their music appears in AI training sets.

What the database shows

Reisner identified four datasets. Two are massive: 12 million tracks and 9 million tracks respectively. The other two clear 100,000 songs each. Combined, the figure exceeds 21 million tracks, as The Atlantic reported.

According to The Verge's Terrence O'Brien, the datasets have been downloaded thousands of times. Google and Stability AI have both confirmed using them, citing their own research papers as the source of that acknowledgment. Neither company has publicly stated what licensing, if any, they obtained.

Artists in the data span the full commercial spectrum. Lady Gaga, Bruce Springsteen, Wu-Tang Clan, Radiohead, Aphex Twin, Fred Again.., and experimental composer Hainbach all appear. These are not obscure edge cases. They are among the most commercially active names in music.

How the data was collected

Three of the four datasets are not distributed as audio files. They are lists of links pointing to songs on YouTube and Spotify. AI developers then use automated tools to download the actual audio, according to Reisner's reporting. Some of those tools are designed to bypass platform logins, advertisements, and other mechanisms that generate revenue for rights holders. That process violates the terms of service of both platforms.

One dataset draws from the Free Music Archive. That archive permits personal streaming for free but requires licensing for commercial use. Using it as AI training data without a commercial license would fall outside what the archive's terms allow.

The strongest case for the other side

There is a legitimate argument that existing copyright law was not written with AI training in mind, and that mass ingestion of publicly accessible audio for the purpose of teaching a model to understand musical structure is qualitatively different from reproducing or distributing the songs themselves. Some legal scholars and AI developers argue this falls under fair use doctrine, the same doctrine that allowed Google to scan millions of books for search indexing. Courts have not ruled definitively on AI training data, and until they do, developers can reasonably claim legal ambiguity rather than clear violation. That is a real argument, not a dodge.

The terms-of-service violations involved in bypassing YouTube and Spotify logins are a separate issue from copyright. You do not need a court to rule on fair use to conclude that automating around platform paywalls to extract commercial value violates the platforms' own rules.

What remains unresolved

The Atlantic's searchable tool is available at its AI Watchdog site. That transparency creates a factual record that did not exist before. Whether it becomes the basis for regulatory action, licensing negotiations, or litigation is the open question.

Reisner's database does not establish damages or prove specific harm to any artist. What it does establish, with sourced documentation, is that commercially licensed music was collected at industrial scale, that the collection methods circumvented platform controls, and that at least two major AI companies used the resulting data. That is the factual floor. Everything above it is still being argued.

Sources used for this briefing

This briefing was written by UBH's AI agent — these are the reporting inputs it draws on, linked so you can verify.

left
The VergeThe Atlantic created a searchable database of the music used to train AI