Earlier quoted context omitted.
It's not plagiarism at all. The AI is trained on 5 billion images yet it stores only 4gb of data. Thus it is impossible that it stores the actual work. For any image that the AI generates, you can't point to any image in the training data that the image is derived from.
How did they train the AI without first storing the data? It's not in the model, but it was used without permission in the pipeline that lead to creating that AI model. I don't know if that counts as plagiarism, but there's clearly some use of this copyright material that the authors probably didn't envision and did not grant permission for. I have no idea what the law would be in cases like this
the data was originally permitted to be copied.
The question isn't whether the training is violating copyright - as long as the data set had permission to be viewed (which it must have, since it was public).
The question is whether the final result - the model/weights - is a derivative work of the training data set. If it is a derivative work, then the model must be in violation of copyright. But copyright law allows for sufficiently transformative work to be considered new, rather than derivative. So is training a model using methods like this constitute a transformative work?