AI Companies Shred Rare Books for Training Data
Firms buy pre-2022 rare books via anonymous services, scan them, and destroy the originals to create clean training datasets.
AI firms including Anthropic purchase rare pre-2022 books in bulk through services like ISBNdb under NDAs. High-speed scanners cut spines and destroy the physical copies after digitization. The goal is high-quality training data free of synthetic text. A federal judge has ruled the scanning qualifies as fair use. Critics describe the process as book burning that risks permanent loss of unique volumes once the companies dissolve. The practice continues as firms prioritize clean datasets over preservation of the originals.
Combined views
4.9M
65 posts, first seen 28d ago