AI · 21h ago
AI firms shred millions of books after using them for training data
AI companies are buying up physical books through middlemen to train large language models, then shredding them, according to a new report. The practice involves secretly acquiring millions of books to build training datasets. This raises questions about copyright and the environmental impact of destroying physical media.
Meridian48 take
The story highlights a wasteful and secretive side of AI data acquisition, but the scale and legality remain unclear.
Read the full reporting
AI companies are reportedly shredding millions of books after using them to train AI models — tech giants outsource to middlemen to secretly buy up books for training material →
Tom's Hardware
ai-trainingcopyright