Earlier quoted context omitted.
> AI models are intellectual property If companies train on data they don't own and expect to own their model weights, that's hypocritical. Model weights shouldn't be copyrightable if the training data was pilfered. But this hasn't been tested because models are locked away in data centers as trade secrets. There's no opportunity to observe or copy them outside of using their outputs as synthetic data. On that subjec…
Now that you mention it, I'm quite surprised that none of the typical fanatical IP lawsuiters had sued arguing (reasonably I think) that the output of the LLMs is strongly suggestive that they have been trained on copyrighted materials. Get the lawsuit to discovery, and those data centers become fair game. Perhaps 'strongly suggestive' isn't enough.
Given that everything -- including this comment -- is copyrighted unless it is (1) old or (2) deliberately put into the public domain, this is almost certainly true.