Earlier quoted context omitted.
> This is a silly opinion to hold, isn't it? I mean, you release projects under a license with the express purpose of freely distributing your code among anyone in the world that may have any interest whatsoever, and even allow they themselves to share it with anyone they feel fit. But you are somehow outraged if people actually use said code? You're making things up: the outrage is not that people used it, it's that…
idk, all the code i've seen produced by an llm doesn't appear to be derived from anything. Also, the source code they were trained on does not exist in the model, it's impossible for the llm to return a code snippet from some other code base. The code snippet doesn't exist in the model in the first place. I guess another way to put it is show your code in the output of an llm that isn't being attributed correctly.
So? Just because a piece of output data is encrypted or compressed and does not resemble the input, does not mean that the process did not take the input.
We have decades of law that regards zipped files as infringment, lossy compression (MP3's) as infringment, etc.
> guess another way to put it is show your code in the output of an llm that isn't being attributed correctly.
Well, a better way of putting it is answering the question "Will that model have existed had none of the code used as input existed".
IOW, can that model be generated or created without first having all that copyrighted code used as input?