it would seem to me that from a technical perspective the weights of an AI model trained on copyrighted material would be a reproduction of the copyrighted work. just because you combine the information from millions (or more) copyrighted works together doesn't mean that you aren't reproducing them. the process of training requires reproduction and distribution of the works internally as part of the data processing p…
Distributing data to servers only accessible by a machine isn't what folks have in mind when they talk about the distribution of copyrighted content.
you would have to ask a judge about that because its a novel legal question. obviously nobody anticipated this technology at the time it was written so ultimately it will have to be a court that decides how to apply existing laws.