Live data from Hacker News

Judge dismisses DMCA copyright claim in GitHub Copilot suit

theregister.com

151–160 of 505 posts

Re: Judge dismisses DMCA copyright claim in GitHub Copilot suit

#151
Big question: this thing called “training” AI off of data, how much of this is “training” and how much of this is “synthesizing”? It seems like if code is being copied and rephrased, it is synthetic. Not much “learning” and “training” going on here.

Re: Judge dismisses DMCA copyright claim in GitHub Copilot suit

#152

Earlier quoted context omitted.

What about the copyrights purpose of furthering the arts and sciences?

Copyright has utterly failed to serve that purpose for a long time, and has been actively counterproductive. But if you want to argue that copyright is counterproductive, I completely agree. That's an argument for reducing or eliminating it across the board, fairly, for everyone; it's not an argument for giving a free pass to AI training while still enforcing it on everyone else .

Without copyright, entire industries would've been dead a long time ago, including many movies, games, books, tv, music, etc.

Just because their lobbies tend to push the boundary of copyright into the absurd doesn't mean these industries aren't worth saving. There should be actually respectful lawmakers who seek for a balance of public and commercial interests.

Re: Judge dismisses DMCA copyright claim in GitHub Copilot suit

#153

A slight aside, but this is the subtitle: > A few devs versus the powerful forces of Redmond – who did you think was going to win? I hate that kind of obnoxious "journalism". Sometimes the little guy is actually wrong. To clarify, I'm not commenting on the specifics of this case, I just hate how fake our online discourse has been by appealing to "big guy evil" before even bringing up the specifics of the case.

It's The Register, they are always like this. Especially when Microsoft is involved.

Re: Judge dismisses DMCA copyright claim in GitHub Copilot suit

#154

Earlier quoted context omitted.

> Point is, the legal system is highly selective when it comes to corporate interests. I don't even think it's that. In recent cases like Oracle v. Google and Corellium v. Apple, Fair Use prevailed with all sorts of conflicting corporate interests at play. The Zenimax v. Oculus case very much revolved around NDAs that Carmack had signed and not the propagation of trade secrets. Where IP is strictly the only thing bei…

In fact, go to far as to argue your example of Authors Guild v. Google is a good indication that most cases will probably go an AI platform's way. It's a pretty parallel case to a number of the arguments. Indexing required ingesting whole works of copyright material verbatim. It utilized that ingested data to produce a new commercial work consisting of output derived from that data. If I remember the case correctly,…

A key finding that the judge said in the Authors Guild v. Google case was that the authors benefited from the tool that google created. A search tool is not a replacement for a book, and are much more likely to generate awareness of the book which in turn should increase sales for the author.

AI platforms that replaces and directly compete with authors can not use the same argument. If anything, those suing AI platforms are more likely to bring up Authors Guild v. Google as a guiding case to determine when to apply fair use.

Re: Judge dismisses DMCA copyright claim in GitHub Copilot suit

#155

> Indeed, last year GitHub was said to have tuned its programming assistant to generate slight variations of ingested training code to prevent its output from being accused of being an exact copy of licensed software. If I, a human, were to: 1. Carefully read and memorize some copyrighted code. 2. Produce new code that is textually identical to that. But in the process of typing it up, I randomly mechanically tweak a…

> How is it any different when a machine does the same thing? Because intent matters in the law. If you intended to reproduce copyrighted code verbatim but tried to hide your activity with a few tweaks, that's a very different thing from using a tool which occasionally reproduces copyrighted code by accident but clearly was not designed for that purpose, and much more often than not outputs transformative works.

It's equally plausible to say you don't intend to reproduce copyrighted code verbatim but occasionally do so given either a sufficiently specific prompt or because the reproduced code is so generic that it probably gets rewritten a hundred times a day because that's how people learned to do basic things from books or documentation or their education.

Re: Judge dismisses DMCA copyright claim in GitHub Copilot suit

#156

Earlier quoted context omitted.

Where it gets ethnically dubious is that: 1. The copilot team rushed to slap a copyright filter on top to keep these verbatim examples from showing up, and now claims they never happen. 2. LLMs are prone to paraphrasing. Just because you filter out verbatim copies doesn't mean there isn't still copyright infringement/plagiarism/whatever you want to call it. The copyright filter is only a legal protection, not a pract…

> The copilot team rushed to slap a copyright filter on top to keep these verbatim examples from showing up, and now claims they never happen. Well if the copyright filter is working they indeed aren't happening. Putting in safe gaurds to prevent something from happening doesn't mean you're guilty of it. Putting a railing on a balcony doesn't imply the balcony with railing is unsafe. > LLMs are prone to paraphrasing.…

> Well if the copyright filter is working they indeed aren't happening. Putting in safe gaurds to prevent something from happening doesn't mean you're guilty of it. Putting a railing on a balcony doesn't imply the balcony with railing is unsafe.

Doesn't mean you weren't, at some point, guilty of it, either. It doesn't retcon things.

Re: Judge dismisses DMCA copyright claim in GitHub Copilot suit

#157

Earlier quoted context omitted.

Ah, you're right. I was wrong to say "public domain". It would be more correct to say Quake III Arena was released to the public as free software under the GPLv2 license.

There is a large gap between public domain and GPL. For starters if Copilot is emitting GPL code for closed source projects... that's copyright infringement.

That would be license infringement, not copyright infringement.

Re: Judge dismisses DMCA copyright claim in GitHub Copilot suit

#158

Earlier quoted context omitted.

I enjoy how you removed the “I think” qualifier which suggested that it’s very possible that you’re right. I’m quite well read on the DMCA but admit you probably know far more about how Nintendo wields it. Still, I suggest that it’s a lot more likely that GitHub is going to get sued than you or GP. Finally, I believe using the legal system to bully independent software developers is, in legal terms, super lame. We ar…

DMCA (at least the take down requests part) is not really suing someone and not really about making money. Its about getting certain works off the internet. You are probably more likely to be on the wrong end of a dmca take down request as a poor person since you dont have the resources to fight it, and its not about recovering damages just censorship.

We are really losing the plot of what this thread is about here, but: DMCA takedown requests that are ignored or wheee the site does not comply with the process are subject to private civil action. Obviously, a takedown request is distinct from suing someone. And the way that the rights holder forces the site to remove the content is under threat of monetary penalties.

Re: Judge dismisses DMCA copyright claim in GitHub Copilot suit

#159
post #4

> The anonymous programmers have repeatedly insisted Copilot could, and would, generate code identical to what they had written themselves, which is a key pillar of their lawsuit since there is an identicality requirement for their DMCA claim. However, Judge Tigar earlier ruled the plaintiffs hadn't actually demonstrated instances of this happening, which prompted a dismissal of the claim with a chance to amend it. I…

Where it gets ethnically dubious is that: 1. The copilot team rushed to slap a copyright filter on top to keep these verbatim examples from showing up, and now claims they never happen. 2. LLMs are prone to paraphrasing. Just because you filter out verbatim copies doesn't mean there isn't still copyright infringement/plagiarism/whatever you want to call it. The copyright filter is only a legal protection, not a pract…

> Just because you filter out verbatim copies doesn't mean there isn't still copyright infringement/plagiarism/whatever you want to call it.

Actually, it does. The production of the output is what matters here.

Re: Judge dismisses DMCA copyright claim in GitHub Copilot suit

#160

Earlier quoted context omitted.

> Someone is likely to design an LLM that is specifically trained to do exactly that. Perplexity AI.

> Perplexity AI. How does this describe Perplexity AI more than any other LLM?

I am referring to their service rather than their LLM in specific.

Perplexity is in the business of using an LLM to paraphrase existing content, then serving that up as their own "work" in a way that directly harms the original content they took.

It's not even a question of "Is AI training copyright infringement", they're just doing copyright infringement with AI. And it's horribly common already.

Post reply on HN