Earlier quoted context omitted.
> It's a neural network. It works similar to our brains work, but more consistent. Irrelevant and incorrect. > It's doesn't seem like piracy to me. It's pretty indisputably piracy, whether or not it's legal/fair use/whatever. Many of the training sets included material like the books3 corpus which was downloaded to a server somewhere. That is simply piracy, doesn't matter why they downloaded it. I believe many artist…
> It's pretty indisputably piracy, whether or not it's legal/fair use/whatever. Ah, this is obviously some strange usage of the word 'indisputably' that I wasn't previously aware of. > I believe many artists rightly refuse to accept this threat to their livelihoods because it was built on their labor . This model is trained from scratch using only public domain/CC0 and copyright images with specific permission for us…
It seems incredible to me to suggest that piracy wasn't involved in the collection of training data, regardless of your view on the morality or legality of it. Datasets like books 3 indisputably contained copyrighted content that was being distributed without permission from the rightsholder. That's just the definition of piracy. If we can't agree on that then I'm not sure what we're doing here.
More materially to this discussion, yes, it would absolutely make a difference if the AI was only trained on licensed content. I wouldn't use it but I wouldn't have a problem with it. The issue is specifically that much of the work being used without permission is being used to replace the people who made that work, and is being used without permission. If the model is based on ethically acquired data, it would be less able to reproduce the style of specific artists. Imo, there would be more room for both kinds of art in this case.
I'm also aware that it's not a clear cut case legally but I think AI advocates and tech enthusiasts think it's a lot more likely that AI will win in court than the actual chances. Napster took years to litigate and was eventually shutdown. There's a really good discussion about this on the decoder podcast between actual lawyers.