Earlier quoted context omitted.
is uncertain, as with codding you need white room methods to prove that new code is not contaminated with patented implementation, as it might be here, so basing anything on an existing model could be also copyrighted.
The model isn't code to a new model trained on it, it's training data; just like the pirated torrent site Books3 dataset Facebook used to train LLaMA. The training code is Apache 2.0 licensed so it can be copied and modified freely, including for commercial purpoes. https://github.com/facebookresearch/llama
But AFAIK this is just the first step to get initial weights and later you need much more work to fine-tune this to get useful results from the model.
I think this step could be seen as contaminating weights with copyrighted content.
Something like chrome is copyrighted but chromium is not
I'm not a lawyer, so I'm not that well informed how official definitions match here, but what's I'm trying to say it that I wouldn't be surprised if this would go either way