TechFlow reports that on July 22, multiple authors filed a class-action lawsuit in 2024, accusing Anthropic of downloading millions of pirated e-books in bulk from "shadow libraries" such as LibGen and incorporating approximately 482,000 copyrighted works among them into Claude's training dataset. The court subsequently separated the determination of "model training" and "data acquisition": AI reading legally obtained books for learning may constitute fair use, but downloading and storing works long-term through pirated websites constitutes infringement in itself. Anthropic ultimately agreed to a $1.5 billion settlement, compensating approximately $3,000 per work on average, and destroying relevant pirated files.
However, Anthropic has previously repeatedly accused some Chinese AI teams of extracting Claude outputs via API for model distillation, claiming it violates intellectual property rules. The two acts are not exactly the same legally but form a stark contrast: when facing human works, AI companies emphasize the model's right to learn; when facing their own model outputs, they demand stronger exclusive protection. Therefore, the core boundary is: AI training is not necessarily illegal, but the method of acquiring training data must be legal. (Jin10)




