When a model is trained on 'all-rights-reserved' content like most image datasets, the community say it's fair game. But when it's 'just-a-few-rights-reserved' content like GPL code, apparently the community says that crosses a line?
Realistically, this tells me that we need ways for people to share things along the lines of all the open-source licenses we see.
You could imagine a GPL-like license being really good for the community/ecosystem: "If you train on this content, you have to release the model."