There are two issues -- (1) feeding copyrighted material in to an AI model, and (2) getting copyrighted material out . The latter is obviously a violation of copyright, full stop. The former, to me, is obviously not a violation. If it were, that would massively tilt the playing field in favor of large corporations. It would become very hard to independently train your own models. Philosophically, I go by the principl…
This was a comment made to me in a previous, similar discussion, discussing case law around Google's use of copyrighted books in building a search engine: https://news.ycombinator.com/item?id=32654478 I'm not sure I completely agree w/ the comment (nor do I think it vindicates CoPilot), but I think it does provide insight into why CoPilot is violating copyright.
The only analogy I can see is that copying the code internally to use in CoPilot training could be a violation of copyright (like how backing up your own MP3s is a violation of copyright?), but the licenses on these public repositories probably already allow that...