Live data from Hacker News

GitHub Copi­lot inves­ti­ga­tion

githubcopilotinvestigation.com

611–620 of 1001 posts

Re: GitHub Copi­lot inves­ti­ga­tion

#611
post #530

Earlier quoted context omitted.

I use Copilot all the time and I’ve never once used it to generate a whole prepackaged function that’s more than maybe three lines. So no, I don’t benefit from its reproducing other people’s code at all. Tell me you don’t use Copilot without telling me about it.

That isn't sufficient to get you off the hook. Copyright covers derivative work, not just code that's reproduced verbatim.

Derivative work requires a major part of the original, before it’s considered by copyright.

A 3 line boilerplate is neither novel nor a major part of the original.

Re: GitHub Copi­lot inves­ti­ga­tion

#612

What do people think the future looks like where publicly available resources on the Internet (art, code, etc) aren't fair use for training ML models? Where you have to opt into models or can opt out (and many wind up doing so)? OpenAI, Microsoft, Google, et al will STILL train such models that can do all the same things, but it will be much harder for non-industry-backed individuals to navigate the legal minefield w…

> that's not how it's going to be used in practice

If I'm writing some code and want the suggestion to be good then why wouldn't I use the name of a top programmer as a prompt?

Re: GitHub Copi­lot inves­ti­ga­tion

#613

Earlier quoted context omitted.

> Here are a few thoughts I haven't formulated before: > It seems clear enough to me that training AIs on copyrighted works is typically or commonly a fair use under existing law, because the AIs can and commonly do learn non-copyrightable elements and aspects of those works. It's very obvious from enormous numbers of examples that current AI systems are capable of learning much more abstract features of human cultur…

AI copyright drama is my favorite gossip these days because it can’t be reconciled until we accept that intelligence is created and held by societies, not individuals. Recent AI is a new way to exercise that intelligence, but it presents a major conflict with capitalism.

Not a conflict with capitalism - it's another trajedy of the commons - the robber barons of old stole the owned commons land and started said 'capitalism'.

Capitalism is still very alive, and will continue to be. It's in conflict with the general welfare of the people...

Re: GitHub Copi­lot inves­ti­ga­tion

#614
post #247

There are two issues -- (1) feeding copyrighted material in to an AI model, and (2) getting copyrighted material out . The latter is obviously a violation of copyright, full stop. The former, to me, is obviously not a violation. If it were, that would massively tilt the playing field in favor of large corporations. It would become very hard to independently train your own models. Philosophically, I go by the principl…

Fair use is baked into copyright law, "full stop". The only way to prevent all uses of your code is to keep it secret. If anyone wants to say me using copilot violates their copyright, then sue me. But if you have no loss of reputation or revenue, and I have an innocent infringer defense - noone can stop me.

> Fair use is baked into copyright law, "full stop".

Someone recently said most statements on HN should automatically get "in the US" appended to them due to how US centric many of the views are. This is an excellent example. There are plenty of juristictions where "fair use" doesn't exist.

Re: GitHub Copi­lot inves­ti­ga­tion

#615
Perhaps the only way out of this is to start suing the users of Copilot, much as some jurisdictions target the users of a product (e.g. drugs, prostitution) as a means to shut it down when the providers are too difficult or numerous to challenge effectively.

Re: GitHub Copi­lot inves­ti­ga­tion

#616

Earlier quoted context omitted.

They’re clearly producing derivative works.

I think you're saying any work created by a model trained on copyrighted data is a derivative work of that copyrighted data. But this can't be right, it is inconsistent with how copyright has worked so far. Artists and musicians and engineers all learn from each other and have seen and learned from, "trained on" many other examples of works from their field. Even when works are clearly inspired by other works we tend…

Yes. The copyright law must change. This is different.

Re: GitHub Copi­lot inves­ti­ga­tion

#617
post #491

It would be sad if someone succeeded in shutting down CoPilot for this kind of copyright stuff. It is genuinely useful. I don't care that it reproduces copyrighted content. The only way you can get it to do that is to bait it with the function names of functions that have already been copy and pasted thousands of times onto GitHub without proper licenses. Luckily, someone will probably come out with a "renegade" vers…

> It is genuinely useful. I don't care that it reproduces copyrighted content.

I feel like the title of the article is literally written for you: """Maybe you don’t mind if GitHub Copi­lot used your open-source code with­out ask­ing. But how will you feel if Copi­lot erases your open-source com­mu­nity?"""

If you want to keep having useful tools based on open source code in the future, it is in your interest that people still want to write open source code. It is still too early to say how much of a chilling effect projects like Copilot will have on that. But clearly many (just read this comment section, myself included) are having second thoughts.

Re: GitHub Copi­lot inves­ti­ga­tion

#618
post #517
post #491

It would be sad if someone succeeded in shutting down CoPilot for this kind of copyright stuff. It is genuinely useful. I don't care that it reproduces copyrighted content. The only way you can get it to do that is to bait it with the function names of functions that have already been copy and pasted thousands of times onto GitHub without proper licenses. Luckily, someone will probably come out with a "renegade" vers…

I call these people open source haters. They selectively choose what they want open source to mean, and are against the fundamental ideas of open source. Long live Copilot. It’s an amazing product that shows what we are capable of thanks to crowdsourcing and bleeding edge technology. We live in the future, and progress never remembers those who tried to stop it.

> I call these people open source haters. They selectively choose what they want open source to mean, and are against the fundamental ideas of open source.

B..but, Copilot isn't open source though?

Re: GitHub Copi­lot inves­ti­ga­tion

#619
post #247

There are two issues -- (1) feeding copyrighted material in to an AI model, and (2) getting copyrighted material out . The latter is obviously a violation of copyright, full stop. The former, to me, is obviously not a violation. If it were, that would massively tilt the playing field in favor of large corporations. It would become very hard to independently train your own models. Philosophically, I go by the principl…

> There are two issues -- (1) feeding copyrighted material in to an AI model, and (2) getting copyrighted material out. > The latter is obviously a violation of copyright, full stop. It's not obvious to me that (2) is a violation of copyright. Unlike patents, copyright violation is not as simple to prove. My understanding is that, at least in the US, independent creation is a valid defense against copyright infringem…

> > There are two issues -- (1) feeding copyrighted material in to an AI model, and (2) getting copyrighted material out.

> > The latter is obviously a violation of copyright, full stop.

> It's not obvious to me that (2) is a violation of copyright. Unlike patents, copyright violation is not as simple to prove. My understanding is that, at least in the US, independent creation is a valid defense against copyright infringement. For example if 2 people independently write the same story and can prove that they did, they can both hold copyright over that story.

> The analogue to this does exist without AI, when creating something that looks like copyright infringement, clean room design (don't look at similar things) is often done to ensure that "independent creation" can be used as a valid defense in court. Given that, I think (1) is probably not safe to do at all if you can't prevent (2).

I don't think the analogue holds, the AI does have direct view of the actual code. In the most paranoid clean room design you have two teams, one analyses the behaviour of some software and writes a specification (without view of the source code), the other then uses that spec to write the reimplementation.

Copilot turns that on its head, you ask to do something it then looks up the source code how to do it and gives that to you.

Re: GitHub Copi­lot inves­ti­ga­tion

#620
post #535

Earlier quoted context omitted.

That's hardly a new thing! For instance, Google search makes billions of dollars by indexing content that other people make.

It should be noted that some juristictions are starting to restrict this (e.g. Australia). Also I would argue if Google would randomly display content of full websites and never post links to the original content it would be in a lot more legal trouble.

> Also I would argue if Google would randomly display content of full websites and never post links to the original content

Google does do this though. Just Google for an easy to answer question, like “when was George Washington born”

Post reply on HN