ATLANTIC REPORTER UNEARTHS 21 MILLION SONGS POWERING AI TRAINING
Lady Gaga, Radiohead, and Bruce Springsteen are in datasets downloaded thousands of times by developers.
by editor5 min readcomments soon

The music industry has spent years wondering exactly what data powers the AI models that can generate a credible pop song on command. Alex Reisner of The Atlantic decided to find out. He uncovered four datasets used to train music AI and made them fully searchable for the public. Two of those datasets are enormous, containing 12 million tracks and 9 million tracks, respectively. The other two are smaller, but still represent a significant amount of training data at over 100,000 songs each. Millions of tracks are freely available in these datasets, even if they aren't supposed to be.
THE SCALE IS UNREAL
The names that pop up range from pop stars like Lady Gaga and Fred Again.. to legends like Radiohead, Aphex Twin, Wu-Tang Clan, and Bruce Springsteen. Experimental composer Hainbach is in there, too. This is not a collection of obscure basement recordings. It is the living and recent history of recorded music, vacuumed up into training sets so AI developers can teach machines to write better hooks.
The sets have been downloaded thousands of times by developers looking to build the next generation of music AI. Google and Stability have both confirmed they used these datasets in research papers. The question of consent is no longer hypothetical. The data is out there, and the biggest names in tech have already fed it into their models.
HOW THE DATA MOVES
How exactly does a dataset end up with 12 million copyrighted tracks without anyone asking permission? The answer is clever and deliberately indirect. According to Reisner, "Three of the datasets I found are distributed as a list of links to songs on YouTube or Spotify."
This is the crucial structural detail. The datasets themselves do not contain the audio files. They are essentially a curated index pointing to where the audio lives on streaming platforms. The actual downloading is handled by the developers using automated tools.
Reisner elaborates on the extraction process: "AI developers download the actual audio using tools that automate the job, some of which allow developers to bypass logins, advertisements, and mechanisms that might earn money or subscribers for creators."
This is where the line between passive scraping and active exploitation gets defined. Bypassing advertisements means bypassing the revenue stream that pays artists for their streams. Bypassing logins means ignoring the wall that separates a public preview from a licensed copy. The tools are doing exactly what they were built to do.
Reisner is direct about the legality of these tools: "Such tools violate the terms of service of these platforms."
Some of the sources, like the Free Music Archive dataset, are free to stream for personal use but require licensing for commercial applications. AI training is indisputably a commercial application. The other datasets do not bother making that distinction at all. They just point to YouTube and Spotify.
THE LEGAL DRAG NET
What Reisner has done is drag the skeleton out of the closet. The AI music generation space has been booming, and companies have been happy to show off the demos. They have been far less transparent about the training data. This investigation provides the most concrete evidence yet of the scale at which copyrighted music is being used without explicit permission.
An honest read of the findings suggests the AI industry built its music capabilities on a mountain of unsanctioned data. Google and Stability are explicitly named as researchers who used these datasets. These are not garage startups. They are two of the highest-value companies in the world. If they are building products on data this contaminated, they are either betting the legal risk is manageable or assuming the rights holders lack the resources to fight 12 million separate battles.
The structure of the datasets also creates a novel legal defence.
If a developer only distributed a list of YouTube links, and someone else downloaded the audio, the developer can argue they never actually possessed the copyrighted recording. The courts will have to decide if facilitating mass downloading from YouTube is the same as prior authorisation, especially when the tools used explicitly violate YouTube's terms of service.
NIGHTMARE FOR ARTISTS
For artists, this is a nightmare scenario. Your music is out there, being used to train the very models that could one day make your job obsolete, and you have almost no way to opt out. Even if you never gave permission, even if your label never signed a deal, if your tracks were on YouTube or Spotify, there is a reasonable chance they ended up in one of these datasets.
The recording industry has been aggressive on copyright in the past, but AI moves faster than lawsuits. By the time the courts figure out the legality of the scraping tools, the models will already be trained. The data has already been downloaded thousands of times. The cat is not just out of the bag. The cat built a machine that can sing like a bird.
THE DEFENSE!
AI developers will point out that web scraping is broadly legal, and that the law has historically treated transformative use as a fair use defence. They will argue that parsing the structure of a song to learn musical patterns is no different from an aspiring musician listening to records to learn chord progressions.
The difference is that the musician cannot memorise a million songs perfectly and regurgitate them on demand without compensation. The law has not always kept up with technology, but it has usually found the line when the financial stakes get high enough.
Reisner's work at The Atlantic is the most thorough public accounting of music AI training data to date. By making it searchable, he has given every artist the ability to check if their work was ingested. The answers will likely force a reckoning that the tech industry has been delaying for years.
what did you make of it?
more from ai
ai
TSMC ADDS $100 BILLION TO ARIZONA CHIP BET, TOTAL HITS $265 BILLION
The additional investment will build at least four more 2nm fabs and advanced packaging, bringing the company's total US commitment to $265 billion.
ai
META WILL ALERT PARENTS IF TEENS DISCUSS SUICIDE WITH META AI
The opt-in feature flags self-harm references in chatbot conversations, with human review before any notification is sent.
ai
ROBLOX'S "BUILD" LETS ANYONE MAKE A GAME FROM THEIR PHONE WITH AI
The new toolset, launching July 28, turns text prompts into playable experiences and puts game creation on iPhone and iPad.
ai
ZOOX REALLS ENTURE ROBOTAXI FLEET OVER SMOKE DETECTION FAILURE
A robotaxi drove into an active fire scene obscured by smoke. NHTSA called emergency scenes not edge cases.





