Anthropic PBC × Various booksellers
In a copyright lawsuit brought by authors, court proceedings in June 2025 revealed that AI company Anthropic had purchased millions of print books, physically disassembled and scanned them, and used the resulting digital files to train its Claude AI models. [30, 32] In a ruling on June 23, 2025, a federal judge in San Francisco determined that this process of buying a physical book, digitizing it, and destroying the original could be considered transformative "fair use" under US copyright law. [28, 30] This was separate from the company's use of pirated digital books, which the court found to be infringing. [30]
The asset
The purchase of millions of physical print books, which were then digitized via "destructive scanning" and used as training data for the Claude large language model.
- 1M books
The record
- Seller status
- Going concern
- Sector
- Book retail
- Modalities
- text
- Jurisdiction
- US
- Docket
- 24-cv-05417
What it means if you hold data
This case highlights a novel and legally distinct method for acquiring training data: the bulk purchase and digitization of physical media. The court's "fair use" ruling on this specific practice, provided the books are legally acquired, creates a potential pathway for AI labs to source high-quality data outside of direct licensing deals with publishers, albeit with significant logistical costs.