Earlier quoted context omitted.
LLMs are being trained on published books a direct equivalent of records. People were able to get one to reproduce over 40% of Harry Potter and the Sorcerer’s stone word for word. https://arstechnica.com/features/2025/06/study-metas-llama-3... There’s zero chance that happened without the book being in their training corpus. Worse, there’s significant effort put into obscuring this.
Yes, they are trained on books, but the courts so far are largely in agreement that AI models are neither copies nor derivative works of the source materials. If you’re searching for legal protections against your works being used for model training, copyright law as written today does not appear to give you cover.
https://www.kron4.com/news/technology-ai/anthropic-copyright...