Live data from Hacker News

We've filed a lawsuit against GitHub Copilot

githubcopilotlitigation.com

751–760 of 824 posts

Re: We've filed a lawsuit against GitHub Copilot

#751

I am not against this lawsuit but I'm against the implications of this because it can lead to disastrous laws. A programmer can read available but not oss licensed code and learn from it. Thats fair use. If a machine does it, is it wrong ? What is the line between copying and machine learning ? Where does overfitting come in ? Today they're filing a lawsuit against copilot. Tomorrow it will be against stable diffusio…

> A programmer can read available but not oss licensed code and learn from it. Thats fair use. If a machine does it, is it wrong ?

My (extremely amateur) understanding is that what is meant by "learn from it" is one of the hinge points of the legal question.

If a programmer reads licensed code and reproduces it verbatim or near-verbatim in a project with a conflicting license, that becomes a legal problem in certain circumstances.

If a programmer reads the same code and gets an idea to implement something different, that's less troublesome (or at least, if it is troublesome it's in a different area; if the idea was related to a patentable process, then other questions arise, but I'm even less qualified to speak to that area of law).

There's nothing special about copy/paste buttons that make them the only way you can infringe copyright.

Fair use doesn't automatically kick in just because someone uses what they took/copied as part of a larger artifact; it's a really complicated legal line.

Re: We've filed a lawsuit against GitHub Copilot

#752
post #714

Earlier quoted context omitted.

You could argue that if the author pursued enforcing their licence over those 66 people their code wouldn't have ended up in the training set in the first place. IANAL but I recall that you can't invoke copyright law to selectively enforce it, copyright is only protected if the holder pursues every violation of it. Maybe it works the same for enforcing a licence.

> I recall that you can't invoke copyright law to selectively enforce it, copyright is only protected if the holder pursues every violation of it IIRC, that is wrong. What you are describing is trademarks , not copyright.

Thanks, I wasn't sure

Re: We've filed a lawsuit against GitHub Copilot

#753

Earlier quoted context omitted.

Moreover, if this case wins, it threatens to disrupt one of the biggest technological progressions of all time. AI/ML will change every field just as the Internet and smartphones did. It doesn't show any indication of peaking, either. If the US chooses the wrong path here, we'll only tie our hands behind our backs. Other countries won't be so foolish. We should be able to train on any media a child could see, hear, o…

> it threatens to disrupt one of the biggest technological progressions of all time. Chill dude, all they have to do is include the licenses on their generated code. If anything, this is going to generate even more progress. The copilot team would have to create some kind of feature that would connect the generated output the the relevant training data. That'd be pretty incredible to see in the field of AI/ML in gene…

If they can actually link output to specific input, the lawsuit has merit and more, GPT-3 is a lie. A neural network is supposed to learn how things work, not memorize a large number of examples and spit them out verbatim - or keep connections to specific inputs.

Copilot losing the lawsuit is evidence it’s a case of overfitting, not true ML.

Re: We've filed a lawsuit against GitHub Copilot

#754
post #506
post #465

Earlier quoted context omitted.

Correct legally, morally, or both? Legally a copyright claim seems weak, but they didn't assert one. Some of their claims look stronger than others. The DMCA claim in particular strikes me as strong-ish at first glance, though. Morally I think this class action is dead wrong. This is how innovation dies. Many of the class members likely do not want to kill Copilot and every future service that operates similarly. Bey…

I'm fine with Copilot, but I think all rightsholders should be allowed to decide if they want their code training it or not. And that should be opt-in, not opt-out. (And refusing to opt in shouldn't have to mean switching to a new hosting platform.) > Beyond that, the class members aren't likely to get much if any money. The only party here who stands to clearly benefit is the attorneys. That's the case in pretty muc…

Hypothetically, if I wanted to learn how to code by studying open source examples on GitHub, should I have to go ask permission of each rightsholder to learn from their code? I agree that, if Copilot is based on a model that overfits to output the exact same code it read, the lawsuit has merit (and Copilot is not really ML), but the idea of ML is that the model doesn’t memorize specific answers, it learns internal rules and higher-level representations that can output a valid result when given some input. Very much like me, the coder, would output valid code when given a use case description, after studying a lot of open source examples of that. Should most programmers just be paying rights to all publishers of code they have studied?

Re: We've filed a lawsuit against GitHub Copilot

#755

Earlier quoted context omitted.

> The person operating the model is not violating the exclusive rights of the copyright author: they are not making copies or derivative works. How do they not make copies? Do you know how a computer works? Ever heard of RAM? (At least the German Urheberrecht recognizes this clearly: You can't do any processing on any data with the help of a computer without at least making temporary local copies , so there are excep…

you are clearly doesn't understand how machine learning works, if machine learning ok copyrighted data becomes illegal then most of our infrastructures will be down because most of it uses machine learning, the first that will affect many people is probably google search

I believe this is the core point of the lawsuit - is Copilot really creating code from what it learned (which happens to, by some weird glitch, mimic the source code) or is it just a big overfitting model that learned to encode and memorize a large number of answers and spit them out verbatim when prompted?

I think that losing this lawsuit has much more serious consequences for Copilot than just having to connect to a list of millions of potential copyright owners - it would mean the model behind it is essentially a failure.

Personal opinion: the real situation lies somewhere in the middle. From what I’ve seen, I think Copilot has some ability to actually generate code, or at least adapt and connect unrelated code pieces it remembers to respond to prompts - but I also believe it just “remembers” (i.e., has a close-to-lossless encoding of the input) how to do some operations and spits them out as part of the response to some prompts.

I hardly think the lawsuit will really explore this discussion, but it sounds like a great investigation into what DL models like transformers actually learn. For all I know, it might even give insight into how we learn. I have no reason to believe that humans don’t use the same strategy of memorising some operations and learning how to adjust them “at the edges” to combine them.

Re: We've filed a lawsuit against GitHub Copilot

#756

Earlier quoted context omitted.

Differentiating between a human and a machine simply because one "is not an algorithm" doesn't make a lot of sense. If it were true, people would very easily game it, by using algorithms to automate the most trivial parts of copying someone's work. Ultimately the algorithm is automating something a human could do. There is a lot of gray area to copyright law, but you can't get around that simply by offloading to an a…

> Differentiating between a human and a machine simply because one "is not an algorithm" doesn't make a lot of sense. Uh? So if I design a self driving car which kills someone, it's the car that goes to jail? Legal precedent seems to indicate this is not the case at all. Because humans and machines are different, simply because humans aren't machines and viceversa.

Last I checked this was very much a gray area. I’d expect at least a long investigation into the amount of work and testing put into validating that the self-driving algorithm operates inside reasonable standards of safe driving. In fact, I expect that, as the industry progresses, the tests for minimal acceptable self-driving safety get more and more standardised.

That doesn’t answer the question of who’s responsible when an accident happens and someone gets hurt or dies - but then, there was a time when animals would be judged and sentenced if they committed a crime under human law. That practice is no longer deemed valid, maybe we need to agree that, if the self-driving car was built with reasonable care, accidents can still happen and it’s no one’s fault.

Re: We've filed a lawsuit against GitHub Copilot

#757

Earlier quoted context omitted.

If a software systematically engages in copyright violation but only haphazardly corrects those violations, those haphazard correct aren't evidence the problem has vanished.

If Copilot is committing widespread infringements of their copyright, then surely they will be able to find examples of such infringement to submit in their lawsuit. I assume they want some kind of broad relief, such as an injunction to take down copilot. They are not going to get it, they are not going to get anything at all, if they can’t even provide examples of violating code.

During the piratebay case, the prosecutor only had to illustrate that it was likely (as in, convinced the judges) that copyright infringement had occurred. They did this by showing the top 100 torrents. They did not have to prove with certainty that the top 100 torrent actually was used by people. The fact that the names of movies and games showed up on the list was enough to convince the judges.

The lawyers defending the founders did try to make the argument that no infringement had been proven, and that the list itself was not proof of any infringement. It was just a list on a website, and they even presented evidence that the counter on the list was algorithm faulty. The judges was not convinced and applied the common sense approach that taken as a whole, it was not believable that no infringement had occurred by the website given the context of the site (the name, the top list, the overall perspective of how the site was designed).

Re: We've filed a lawsuit against GitHub Copilot

#758
post #427

Earlier quoted context omitted.

Maybe researchers that are used to hunting for publications and attributions. If I’m sharing my code publicly, it’s because I want it to be _used_.

People share proprietary code publicly. And the fact that you're allowed to read a book doesn't (currently) give you the right to copy it and redistribute the copy.

If I read 10 or 20 books about a topic and then go teach that topic to others, do I have to attribute each thing to all the authors from where I learned it? And what if I come up with my own interpretation of a topic, do I have to trace it back to all interpretations of all the authors that influenced it? Even more, do the previous authors also have to do that and do I have to quote all the chain of references? If not, why an ML model that is supposed to learn how coding works, not memorize pieces of code verbatim, should have to “because of copyright laws”?

Re: We've filed a lawsuit against GitHub Copilot

#759
post #306

Earlier quoted context omitted.

Attributions are fundamental to open source? I thought having source openly available was fundamental to open source (and allowed use without liability/warranty) as per apache, mit, and other licenses. If they just stick to using permissive-licensed source code then i'm not sure what the actual 'harm' is with co-pilot. If they auto-generate an acknowledgement file for all source repos used in co-pilot, and then asked…

Attribution and inclusion of copies of licenses are stipulations in almost all of the popular open source licenses, including BSD and MIT licenses.

True and valid. But all those clauses, AFAIK, were written with the mindset of ïf you want to run this code (particularly, but not limited to, for profit), you have to at least attribute it”. Copilot allegedly doesn’t run that code - it claims to read it, understand how it works and then generate its own code that does an equivalent function if requested. It’s up to the lawsuit to decide if that’s what it actually does, but my point is that the licenses simply did not cover this usage pattern, as much as no open source license requires any kind of action from someone who’s just reading or studying the code.

Re: We've filed a lawsuit against GitHub Copilot

#760
post #329

Earlier quoted context omitted.

> “AI” is just fancy speak for “complex math program” Not really? It's less about arithmetic and more about inferencing data in higher dimensions than we can understand. Comparing it to traditional computation is a trap, same as treating it like a human mind. They've very different, under the surface. IMO, if this is a data problem then we should treat it like one. Simple fix - find a legal basis for which licenses a…

Who decides what constitutes an "AI program" vs just a "program"? What heuristic do we look at? At the end of the day, they have an equivalent of a .exe which runs, and outputs code that has a license attached to it.

I can suggest an idea, considering that the “AI program” is the model, not the training algorithm.

A program gets written by an entity (usually a person) and is executed to generate the desired output according to a deterministic mathematical function it expresses. A training algorithm is a program that gets written to train a model (the model being the “AI Program”) when presented to some training data inputs, to implement a function that is not the training algorithm function itself, but another one, generalising over a problem domain beyond just the original examples fed to the training algorithm.

The output model is not the training algorithm or the training data (or an encoding of it) and exists as its own artefact, independent of both.

Post reply on HN