Live data from Hacker News

GitHub Copilot, with “public code” blocked, emits my copyrighted code

twitter.com

411–420 of 806 posts

Re: GitHub Copilot, with “public code” blocked, emits my copyrighted code

#411

Howdy, folks. Ryan here from the GitHub Copilot product team. I don’t know how the original poster’s machine was set-up, but I’m gonna throw out a few theories about what could be happening. If similar code is open in your VS Code project, Copilot can draw context from those adjacent files. This can make it appear that the public model was trained on your private code, when in fact the context is drawn from local fil…

> This is a new area of development, and we’re all learning. I’m personally spending a lot of time chatting with developers, copyright experts, and community stakeholders to understand the most responsible way to leverage LLMs. Given that there have been major concerns about copyright infringements and license violations since the announcement of Copilot, wouldn't it have been better to do some more of this "learning…

> why not train it on opt-in repositories for a few years first, and iron out the kinks?

Ha ha. Because then the product couldn’t be built. Better to steal now and ask forgiveness later, or better yet, deny the theft ever occurred.

Re: GitHub Copilot, with “public code” blocked, emits my copyrighted code

#412
The "AI" that people keep talking about is no different than any other app like MS Word, which is just a piece of software that serve corporation interests. What we are experiencing today is very simple - big players are using people's work for profit without paying one cent or getting any consent, no need to talk about "How". This is a nightmare scenario under today's social and eco system, and even worse at a production level, because in the end it will form a new industry that has nothing to do with experienced people. Take creative work for example, at the current rate most artists will completely decouple from industry in several years while giving all their works for training for free, and those who control the H/W/R&D resources will find ways to profit from model one way or another, resulting in an "AI" companies controlled "creative industry" with few artists left to direct their work. Can't even think of any other examples close to this in modern history, that a small group of people can do whatever they want under the disguise of "Exciting Technology" which in reality is just stealing an entire industry. There's very little to discuss if you ignore the reality of social systems and just focusing on technical details. We don't live in some fairy tale where you can just let computer do your work and enjoy your life.

Re: GitHub Copilot, with “public code” blocked, emits my copyrighted code

#413

Earlier quoted context omitted.

What makes them bad? I am against Copilot because Microsoft is training the models with public data disregarding copyright (also, doesn't include it's own code).

Not the OP but I have a sinking feeling that these AI tools are going to take away from the most enjoyable careers and creative pursuits and leave us with only mundane button pusher AI supervisor jobs. Current AI is not replacing anything yet but I feel we are only a few years before AI can do a better job at drawing or programming than someone with years of practice. Sure, you can utilise those tools to stay ahead.…

It seems like these AI tools, if anything, will take away the least enjoyable parts of creative careers. Artists will less time thinking about how to adjust a camera lens or mix paints, and more time thinking about how to tell a story that connects with people. Programmers will spend less time banging out boilerplate, and more time thinking about system design.

Re: GitHub Copilot, with “public code” blocked, emits my copyrighted code

#414

I’ve noticed that people tend to disapprove of AI trained on their profession’s data, but are usually indifferent or positive about other applications of AI. For example, I know artists who are vehemently against DALL-E, Stable Diffusion, etc. and regard it as stealing, but they view Copilot and GPT-3 as merely useful tools. I also know software devs who are extremely excited about AI art and GPT-3 but are outraged b…

The last time I happened to point this out[1], all I got was a bunch of HNers nitpicking the words I chose, but not addressing the core issue.

I have to assume this is just people being protective of their own profession and consequently, setting up a high bar for what constitutes as performance in that profession.

[1] https://news.ycombinator.com/item?id=32895251#32895709

Re: GitHub Copilot, with “public code” blocked, emits my copyrighted code

#415

Earlier quoted context omitted.

I'm not even a good cellist and YouTube has put copyright claims on the crappy practice videos I have of me playing Saint-Saëns.

I suspect a video of you playing literally anything on the cello--even an improvised song or a random motif--is likely to get reported as a copyright violation when uploaded to YouTube.

Interesting theory. I'll have to test that.

Oh, actually I remember now -- I think the copyright complaint specifically said what recording they thought I was infringing, and it was the correct piece.

Re: GitHub Copilot, with “public code” blocked, emits my copyrighted code

#416
post #346

Howdy, folks. Ryan here from the GitHub Copilot product team. I don’t know how the original poster’s machine was set-up, but I’m gonna throw out a few theories about what could be happening. If similar code is open in your VS Code project, Copilot can draw context from those adjacent files. This can make it appear that the public model was trained on your private code, when in fact the context is drawn from local fil…

This doesn’t at all address the primary issue, which is one of licensing. Is it a valid defense against copyright infringement to say “we don’t know where we got it, maybe someone else copied it from you first?” If someone violated the copyright of a song by sampling too much of it and released it in the public domain (or failed to claim it at all), and you take the entire sample from them, would that hold up in a le…

> Is it a valid defense against copyright infringement to say “we don’t know where we got it, maybe someone else copied it from you first?”

I mean, in humans it's just referred to as 'experience', 'training', or 'creativity'. Unless your experience is job-only, all the code you write is based on some source you can't attribute combined with your own mental routine of "i've been given this problem and need to emit code to solve it". In fact, you might regularly violate copyright every time you write the same 3 lines of code that solve some common language workaround or problem. Maybe the solution is CoPilot accompanying each generation with a URL containing all of the run's weights and traces so that a court can unlock the URL upon court order to investigate copyright infringement.

> If someone violated the copyright of a song by sampling too much of it and released it in the public domain (or failed to claim it at all), and you take the entire sample from them, would that hold up in a legal setting? I doubt it.

In general you're not liable for this. While you still will likely have to go to court with the original copyright holder's work, all the damages you pay can be attributed to whoever defrauded or misrepresented ownership over that work. (I am not your lawyer)

Re: GitHub Copilot, with “public code” blocked, emits my copyrighted code

#417

Earlier quoted context omitted.

I often quote this comment regarding AI advances and jobs [0]: > Yes, many of us will turn into cowards when automation starts to touch our work, but that would not prove this sentiment incorrect - only that we're cowards. >> Dude. What the hell kind of anti-life philosophy are you subscribing to that calls "being unhappy about people trying to automate an entire field of human behavior" being a "coward". Geez. >>> B…

It’s not hypocrisy to think some jobs shouldn’t be automated. I don’t teach, but I definitely want human teachers teaching my kin, not AI teachers.

Perhaps, or perhaps not. We have not yet seen the true reach of pedagogy of AI. If AI can teach better than humans (something like the Matrix's brain uploading of training), then I will want to do that than have a human teach me.

Re: GitHub Copilot, with “public code” blocked, emits my copyrighted code

#418
post #221

Other than the legality of this code being copied almost verbatim - who is the person who would use it? A person who would use it could also write it. If they cannot write it, why would they ever use this very specific code that magically appeared?

for the same reason people copy random snippets from stackoverflow without understanding how they work. There's a large amount of people who care more about getting a job done than really understanding how the tools work they're using. There is absolutely zero doubt in my mind that copilot et al will lead to the absolute proliferation of half baked code even more than all the other mundane ways to copy&paste do.

> There is absolutely zero doubt in my mind that copilot et al will lead to the absolute proliferation of half baked code even more than all the other mundane ways to copy&paste do.

Agreed. That was my point.

Re: GitHub Copilot, with “public code” blocked, emits my copyrighted code

#419
post #208

Earlier quoted context omitted.

The code is not GPL but is copyrighted in his name.

GPL'd code has a copyright owner. those two things exist at the same time. try reading a licence now and again!

Slow down there with the snarky comments. I never said GPL'd code doesn't have a copyright owner.

Re: GitHub Copilot, with “public code” blocked, emits my copyrighted code

#420

Earlier quoted context omitted.

Uhh, I'm gonna have to disagree hard on this take: > So obviously the source of the error is one or more third parties, not Microsoft, and it's obviously impossible for Microsoft to be responsible in advance for what other people claim to license. Copilot is Github's product, and Microsoft owns Github. They are responsible for how that product functions. In this case, they should be held responsible for the training…

> Giving them the benefit of the doubt here, at minimum they chose which random third parties to believe were honest and correct. Well probably no, they didn't pick and choose at all, they just "chose" everyone who put code online with a license. Which is a legal statement of ownership by each of those people, and implies legal liability as well. > is that the fault of Stephen King or the person selling the robot? We…

> Well probably no, they didn't pick and choose at all, they just "chose" everyone who put code online with a license. Which is a legal statement of ownership by each of those people, and implies legal liability as well.

What you're describing is a choice. They chose which people to believe, with zero vetting.

> The point is that with ML training data, such a vast quantity is required that it's unreasonable to expect humans to be able to research and guarantee the legal provenance of it all.

I'm not sure what you're presenting here is actually true. A key part of ML training is the training part. Other domains require a pass/fail classification of the model's output (see image identification, speech recognition, etc.) so why is source code any different? The idea that "it's too much data" is absolutely a cop-out and absurd, especially for a company sitting on ~$100B in cash reserves.

Your argument kind of demonstrates the underlying point here: They took the cheapest/easiest option and it's harmed the product.

> A crawler simply believes that licenses, which are legally binding statements, are made by actual owners, rather than being fraud. It does seem reasonable to address the issue with takedowns, however.

Yes, and to reiterate, they chose this method. They were not obligated to do this, they were not forced to pick this way of doing things, and given the complete lack of transparency it's a large leap of faith to assume that their training data simply looked at LICENSE files to determine which licenses were present.

For what it's worth, it doesn't seem that that's what OpenAI did when they trained the model initially in their paper[1]:

    Our training dataset was collected in May 2020 from 54 mil-
    lion public software repositories hosted on GitHub, contain-
    ing 179 GB of unique Python files under 1 MB. We filtered
    out files which were likely auto-generated, had average line
    length greater than 100, had maximum line length greater
    than 1000, or contained a small percentage of alphanumeric
    characters. After filtering, our final dataset totaled 159 GB.
I have not seen anything concrete about any further training after that, largely because it isn't transparent.

[1]: https://arxiv.org/pdf/2107.03374.pdf

Post reply on HN