Earlier quoted context omitted.
> To me it sounds like the AI-written work can not be coppywritten I think we didn't even began to consider all the implications of this, and while people ran with that one case where someone couldn't copyright a generated image, it's not that easy for code. I think there needs to be way more litigation before we can confidently say it's settled. If "generated" code is not copyrightable, where do draw the line on wha…
All of your questions have seemingly trivial answers. Maybe I am missing something, but... > If "generated" code is not copyrightable, where do draw the line on what generated means? Do macros count? Does the output of the macro depend on ingesting someone else's code? > Does code that generates other code count? Does the output of the code depend on ingesting someone else's code? > Protobuf? Does your protobuf imple…
> the complaint is not about code generation, it's about ingesting someone else's code, frequently for profit.
Why do you think that is, and what complaint specifically? I was talking about this:
> The Copyright Office reviewed the decision in 2022 and determined that the image doesn't include “human authorship,” disqualifying it from copyright protection
There seems to be 0 mentioning of training there. In fact if you read the appeal's court case [1] they don't mention training either:
> We affirm the denial of Dr. Thaler’s copyright application. The Creativity Machine cannot be the recognized author of a copyrighted work because the Copyright Act of 1976 requires all eligible work to be authored in the first instance by a human being. Given that holding, we need not address the Copyright Office’s argument that the Constitution itself requires human authorship of all copyrighted material. Nor do we reach Dr. Thaler’s argument that he is the work’s author by virtue of making and using the Creativity Machine because that argument was waived before the agency.
I have no idea where you got the idea that this was about training data. Neither the copyright office nor the appeals court even mention this.
But anyway, since we're here, let's entertain this. So you're saying that training data is the differentiator. OK. So in that case, would training on "your own data" make this ok with you? Would training on "synthetic" data be ok? Would a model that sees no "proprietary" code be ok? Would a hypothetical model trained just on RL with nothing but a compiler and endless compute be ok?
The courts seem to hint that "human authorship" is still required. I see no end to the "... but what about x", as I stated in my first comment. I was honestly asking those questions, because the crux of the case here rests on "human authorship of the piece to be copyrighted", not on anything prior.
[1] - https://fingfx.thomsonreuters.com/gfx/legaldocs/egpblokwqpq/...