Earlier quoted context omitted.
I think that’s a more interesting question. I’m not sure how to do it, but I think that finding a way to stop the reproduction of copyrighted content is probably the missing piece. If there was a monetary penalty for reproduction of copyright works (apply human laws to machine), then I bet these companies would quickly figure out how to fingerprint output and match it to source data before sending it to the user.
So I go into a cinema with my recording equipment, record the movie, step outside and start selling downloads, and when the cops come, I say: "I’m not sure how to do it, but I think that finding a way to stop the reproduction of copyrighted content is probably the missing piece." Really?
If LLMs were specially advertised as a way to get all the stuff you already love for free, it would be.
They're not.
What they are advertised as is a way to solve problems and create novel things.
In this regard, LLMs are less of a copyright infringement issue than, e.g. Google News and Google Images both of which had to change because they were in law copyright infringements.
This does not mean that LLMs or diffusion models must get a free pass or anything like that: these are a novel things that didn't previously exist, so while I was initially surprised by the legal cases against Stability and OpenAI due to the existence of Google search and that nobody seemed to care about GPT-3 or the original DALL•E, the arguments made against them are nevertheless interesting and worth caring about.
I suspect LLMs are so useful the powers that be will just carve out a space for them; conversely I don't see that applying to image generation models, so I kinda expect them to be strictly limited to cases where the source data can be proven correctly licensed.