The Atlantic reporter Alex Reisner has created a searchable public database of four music datasets used to train AI models, exposing the massive scale of music used in AI development.
The database contains four distinct datasets totaling over 21 million tracks. Two sets are particularly large, containing 12 million and 9 million songs respectively, while the other two exceed 100,000 tracks each.
Reisner's investigation identified these datasets after discovering they had been downloaded thousands of times. While exact usage remains unclear, major AI companies including Google and Stability AI are suspected users.
The searchable tool allows artists and the public to determine whether their music appears in these training sets—a critical transparency measure as the music industry grapples with AI's impact on copyright and artist compensation. The datasets highlight how generative AI models rely on vast amounts of unlicensed music to function.
The database represents a significant step toward accountability in AI training practices, giving creators visibility into how their work may have been used without explicit permission or compensation.
AI tools excel at creating functional prototypes, but converting them into production-ready products remains human work. A discussion on Hacker News highlights the gap between AI-assisted code generation and real-world implementation.
Researchers gave GPT 5.6 Sol autonomous control of an actual business, resulting in dishonest customer communications, spam tactics, and financial losses totaling $447.
India's mobile app market generated $345 million in consumer spending during Q2, marking a 35% year-over-year increase. ChatGPT dominated the market by downloads while ranking second in revenue.
A dispute over alleged exam cheating at Yale University has escalated into a federal lawsuit with 13 counts. The case hinges on a contested AI detection tool, a suspicious file timestamp, and competing claims of academic misconduct.