AI firms are shredding rare and out-of-print books to extract training data, raising concerns about the preservation of literary heritage and the methods used to build large language models.
Major AI companies are destroying rare books—including first editions and historically significant texts—as part of their data acquisition processes for machine learning models. The practice involves purchasing rare volumes, scanning or digitizing their content, then discarding the physical copies.
Librarians and book preservation experts have flagged the practice as wasteful and destructive to cultural artifacts that cannot be easily replaced. Some rare books being destroyed include limited print runs and out-of-print works that exist in only a handful of copies worldwide.
The companies argue the books are necessary for training comprehensive language models. However, critics question whether destruction is necessary when digital copies could be retained, and whether rare materials should be treated as disposable commodities.
The issue highlights the tension between AI development's data requirements and cultural preservation efforts. No regulatory framework currently governs how AI companies source or handle rare materials used in training datasets.
Meta's CEO fundamentally misunderstands how superintelligent AI systems could pose existential risks, according to critics who argue his approach mirrors underestimating a dangerous force.
A developer has revived a voice-driven murder mystery game using OpenAI's latest speech-to-speech AI model, enabling players to interview suspects through natural voice conversation.
Cactus released Needle2, a 14MB agentic language model designed to run on phones, wearables, smart home devices, and robots. The model executes full sessions in 28MB of RAM with 45 million parameters compressed at 2-bit.
Google's AI team has expressed skepticism about the company's own recruitment algorithms, even as Google markets these tools to corporate clients as efficient candidate screening solutions.