The Atlantic created a searchable database of the music used to train AI

https://platform.theverge.com/wp-content/uploads/sites/2/2026/01/STK467_AI_MUSIC_CVirginia_B.jpg?quality=90&strip=all&crop=0%2C10.732984293194%2C100%2C78.534031413613&w=1200

Millions of tracks are freely available in datasets, even if they’re not supposed to be.

by Terrence O'Brien

Jun 20, 2026, 6:46 PM UTC

Image: Cath Virginia / The Verge

Terrence O'Brien is the Verge’s weekend editor. He’s covered the tech industry for over 18 years and knows a thing or two about synths.

Atlantic reporter Alex Reisner recently uncovered four datasets of music being used to train AI models and made them fully searchable for the public. Two of the sets are absolutely enormous at 12 million and 9 million tracks. The other two are much smaller, but still represent a significant amount of training data at over 100,000 songs each.

According to Reisner, the sets have been downloaded thousands of times and, while it’s impossible to know exactly who has used them, Google and Stabilityhave both confirmed they have in research papers. Some of the sources, like...

Copyright of this story solely belongs to theverge.com. To see the full text click HERE

Read more