The entire discussion about AI and copyright strikes me as a bit naive.
Right now, we are in a situation where nobody quite knows what these AI models are useful for. We have some inkling that they might be extraordinarily useful for making money -- but not precisely how, not even the companies that are developing the models themselves.
Once they money starts, the debate over copyright will fall exactly into the economic seams between the major players involved:
- new tech orgs who are monetizing models will say that the model is "exactly as humans are": they see copyrighted works in training, and then produce wholly original outputs. And of course that the model weights themselves are, like the outputs of employees, completely owned by the company.
- incumbents who stand to lose out on the new gold rush will say that every single output of a model belongs to them if just a single image or sentence was seen in training. And that because of that, we really should just shut the whole thing down, because how could you ever prove that a model was not trained on copyrighted material?
The faultlines will entirely rest on who has more power, hard and soft. How much can they influence the legal system, either by spending $ to hire legal talent or by sheer soft politicking, balanced with how favorable they appear to the general public who uses their product (or consumes their media). I suspect that the end result of this debate is a "legal" way of doing things accessible only to the extremely large players, and a small, politically insignificant collection of individuals, hackers, and startups who aim to unseat those large players (or just flat-out train "illegal" models). The worst possible end result is that the legal system is just too fossilized to deal and tries something draconian like not allow datacenter-scale GPU compute.
As an aside, I predict a sizeable space for companies that do "compliance" -- asserting the copyright status of a dataset, perhaps even themselves using ML. That market will carve off and leave rotting a sizeable chunk of the new money's ML profits.
It's fun to talk about this, I guess. But remember that what you or I have to say about what a machine learning model philosophically is has no bearing what-ever when it comes to the actual ability for individuals, startups, or large players to use models.
I will predict though: enjoy Llama2 while it lasts. Like the internet, it will become fully assimilated into the larger intellectual property machine.