READ. SCROLL. LISTEN.

Original briefings. Zero spin.

Every story is an original briefing written from 60+ sources across the spectrum — sources linked so you can verify it yourself.

← Back to headlines

Atlantic's AI Music Database Draws Scrutiny Over Terms-of-Service Violations in How Datasets Were Built

Atlantic's AI Music Database Draws Scrutiny Over Terms-of-Service Violations in How Datasets Were Built
Since The Atlantic's AI music database went public, the spotlight has shifted from which songs appear in AI training sets to exactly how those songs got there. Three of the four datasets were assembled using tools that automate downloads from YouTube and Spotify in ways that violate both platforms' terms of service.

Since The Atlantic's searchable AI music database became public, the conversation has moved past the headline names — Lady Gaga, Bruce Springsteen, Wu-Tang Clan, Radiohead — and into something more specific: how did millions of copyrighted tracks end up in AI training sets in the first place, and does the method matter legally.

How the Data Was Actually Collected

Atlantic reporter Alex Reisner identified four datasets used to train AI music models. Two are massive: 12 million tracks and 9 million tracks respectively. Two others clock in at over 100,000 songs each. According to Reisner's reporting in The Atlantic, three of the four datasets were NOT distributed as audio files. They were distributed as lists of links pointing to songs hosted on YouTube and Spotify.

AI developers then used automated tools to download the actual audio from those links. Reisner writes that some of those tools are specifically designed to bypass logins, advertisements, and monetization mechanisms. That means creators earned nothing — not streaming revenue, not ad revenue, not subscription credit — when their music was pulled for training data.

Those download tools violate the terms of service of both YouTube and Spotify, according to Reisner's reporting.

Who Used Them

The datasets have been downloaded thousands of times, according to The Atlantic. Pinning down exactly who used them for what purpose is genuinely difficult — these are public research repositories, not internal corporate systems.

Two companies, however, confirmed use in published research papers: Google and Stability AI. Neither company disputes that confirmation, according to the reporting. What remains unknown is whether other major AI music platforms drew on the same sources. Suno recently attracted $400 million in fresh investment, according to The Verge, which reported that figure separately on June 4.

The Licensing Gap

One of the datasets, the Free Music Archive collection, is particularly instructive. The tracks in it are freely available for personal streaming. Commercial use requires a license. Training a commercial AI model almost certainly qualifies as commercial use — though that question remains legally contested.

Developers have argued, in other copyright contexts, that training on publicly accessible data constitutes fair use. The strongest version of that argument: transformative AI models don't reproduce individual songs, they learn statistical patterns from them, which is categorically different from copying. Several ongoing lawsuits in other domains (visual art, text) are testing that theory, but music-specific litigation at this scale has not yet produced a controlling precedent.

The Strongest Counterargument Deserves a Fair Hearing

The fair use case for AI training is not frivolous. Academic researchers have scraped publicly accessible web content for decades to build models, corpora, and indexes, and courts have generally permitted it. If the legal standard for AI training data is "publicly accessible," then a list of YouTube links technically points to public content. The developers who used these datasets may have genuinely believed, and may still believe, they were operating within the law.

That argument weakens, though, when the download tools specifically circumvent monetization mechanisms. Bypassing an ad or a login wall is not just a terms-of-service technicality. It is a deliberate step that denies the creator the compensation the platform is designed to deliver. Whether that rises to legal liability is unsettled. Whether it is ethical is a different question, and on that one the facts are harder to spin.

What the Database Actually Shows

The Atlantic's AI Watchdog tool, now publicly searchable, lets anyone look up whether their music, writing, or other creative work appears in these training sets. The Verge confirmed the tool covers songs, books, and other media categories, not just music.

The database itself does not establish infringement. What it establishes is a factual record: these specific tracks were present in datasets that were downloaded, in some cases using tools that violated platform terms, and were then used — confirmed in at least two cases — to train commercial AI systems.

The unresolved question that matters most going forward: whether the terms-of-service violations in the data collection process affect the legal status of the AI models built from that data. If a court finds that the method of collection taints the training set, that could expose Google, Stability AI, and any other confirmed user to damages that dwarf the licensing fees they avoided.

Sources used for this briefing

This briefing was written by UBH's AI agent — these are the reporting inputs it draws on, linked so you can verify.

left
The VergeThe Atlantic created a searchable database of the music used to train AI
left
The AtlanticThe Atlantic: The Hidden Data Used to Train AI Music Generators