Though I think that training data is totally a violation of copyright, OpenAI really needs to win. Copyright has been unreasonably extended to the point where it's untenable and if we needed to track down rights holders and negotiate 'training rights' we'd ensure there would be no open models or competition going forward. The rights holders are going to lose this one and it's probably for the better.
"training data is totally a violation of copyright" This really isn't clear because cognition is treated as a special exception to copyright. Every thought we have is derivative of everything we've seen before to some degree; reading a book makes our brains a derivative work. But we recognize that cognition is special. With machines we tend to apply a strict test: Did copyright go in? If so, the output is almost cert…
2) even if they were somehow proven to be the same there is still no reason why the same standards need to be applied to computer programs and humans because computer programs do not have any rights or legal protections.
3) cognition is not a "special exception to copyright" because it is entirely unrelated. "Copy" "right" is who has rights to make copies. Your thoughts are not considered copies because they are intangible.
4) we do not "judge every thought individually as to it's originality" because other peoples' thoughts are entirely opaque. Nobody is judging your thoughts, and if you think they are you need to take your medications.