Earlier quoted context omitted.
Exactly. A lot of the difficulty here is how they skip is the hugely important issue: An entirely reasonable, if not fully tested, statement is the following: Every single one of these AI weight things itself is a result of unencumbered, massive, law-breaking, right-violating copyright infringement -- accordingly, it's extremely difficult to say anything morally justifiable or authoritative about anyone elses "rights…
> is a result of unencumbered, massive, law-breaking, right-violating copyright infringement Is there an official ruling? Or is it just a Reddit-style over exaggeration?
They are trained on a lot of text. News sites, comments, books etc. Most books and news sites fall under copyright. Is this fair use? Who knows. Fair use is also an American thing. ChatGPT can be used in the EU, which doesn't have such a broad view of fair use.
If you make a game only out of a lot of copyrighted assets without paying it isn't fair use. Are LLMs different?
What about image generation, which you can prompt the models for specific styles of artists, which works are all copyrighted, but still used for training?