Live data from Hacker News

Judge dismisses DMCA copyright claim in GitHub Copilot suit

theregister.com

171–180 of 505 posts

Re: Judge dismisses DMCA copyright claim in GitHub Copilot suit

#171

> Indeed, last year GitHub was said to have tuned its programming assistant to generate slight variations of ingested training code to prevent its output from being accused of being an exact copy of licensed software. If I, a human, were to: 1. Carefully read and memorize some copyrighted code. 2. Produce new code that is textually identical to that. But in the process of typing it up, I randomly mechanically tweak a…

> How is it any different when a machine does the same thing? Because intent matters in the law. If you intended to reproduce copyrighted code verbatim but tried to hide your activity with a few tweaks, that's a very different thing from using a tool which occasionally reproduces copyrighted code by accident but clearly was not designed for that purpose, and much more often than not outputs transformative works.

> clearly was not designed for that purpose,

I'm not aware of evidence that support that claim. If I ask ChatGPT "Give me a recipe for squirrel lemon stew" and it so happens that one person did write a recipe for that exact thing on the Internet, then I would expect that the most accurate, truthful response would be that exact recipe. Anything else would essentially be hallucination.

Re: Judge dismisses DMCA copyright claim in GitHub Copilot suit

#172
post #3

This is pretty interesting, and I have conflicted feelings about the (seemingly obvious) outcome of this trial. I wonder, if MS and OpenAI win, does that mean it will be legal for anyone to take the leaked source code for a proprietary product, train an LLM on it, and then ask the LLM to emit a version of it that is different enough to avoid copyright infringement? That would be quite the double-edged sword for propr…

I suspect that this is exactly what will happen; not just with code, but also prose and artwork. Someone is likely to design an LLM that is specifically trained to do exactly that. Lots of money to be made...

I was mainly inspired by this section:

> Specifically, the judge cited the study's observation that Copilot reportedly "rarely emits memorized code in benign situations, and most memorization occurs only when the model has been prompted with long code excerpts that are very similar to the training data."

That almost sounds like it'd be fine to train an "art transformation model" which takes an image and transforms it, which for all the frames of a specific Disney movie just so happen to output the very next frame...

Re: Judge dismisses DMCA copyright claim in GitHub Copilot suit

#173

Earlier quoted context omitted.

If the person didn't have your permission or permission from the license to agree to github's terms, then you sue the person who uploaded it to GitHub. You don't get to go after GitHub because you have no contractual relationship with them. At best, you can get an injunction forcing them to take it down, though getting them to un-train copilot may not be feasible. At best you'd get a small cash offer, since you're un…

17 USC §504 says otherwise: ... the copyright owner may elect, at any time before final judgment is rendered, to recover, instead of actual damages and profits, an award of statutory damages for all infringements ... in a sum of not less than $750 or more than $30,000. ... in a case where the copyright owner sustains the burden of proving, and the court finds, that infringement was committed willfully, the court in i…

[deleted]

Re: Judge dismisses DMCA copyright claim in GitHub Copilot suit

#174

> Indeed, last year GitHub was said to have tuned its programming assistant to generate slight variations of ingested training code to prevent its output from being accused of being an exact copy of licensed software. If I, a human, were to: 1. Carefully read and memorize some copyrighted code. 2. Produce new code that is textually identical to that. But in the process of typing it up, I randomly mechanically tweak a…

The actual answer here, regardless of a court ruling, is that you'd go broke if anyone big enough tried to go after you for it.

Legal protections for source code are still pretty fuzzy, understandably so given how comparatively new the industry is. That doesn't stop lawyers from racking up huge fees though, it actually helps because they need so much more prep time to debate a case that is so unclear and/or lacking precedent.

Re: Judge dismisses DMCA copyright claim in GitHub Copilot suit

#175

Earlier quoted context omitted.

> Point is, the legal system is highly selective when it comes to corporate interests. I don't even think it's that. In recent cases like Oracle v. Google and Corellium v. Apple, Fair Use prevailed with all sorts of conflicting corporate interests at play. The Zenimax v. Oculus case very much revolved around NDAs that Carmack had signed and not the propagation of trade secrets. Where IP is strictly the only thing bei…

In fact, go to far as to argue your example of Authors Guild v. Google is a good indication that most cases will probably go an AI platform's way. It's a pretty parallel case to a number of the arguments. Indexing required ingesting whole works of copyright material verbatim. It utilized that ingested data to produce a new commercial work consisting of output derived from that data. If I remember the case correctly,…

> In fact, go to far as to argue your example of Authors Guild v. Google is a good indication that most cases will probably go an AI platform's way.

The more recent Warhol decision argues quite strongly in the opposite direction. It fronts market impact as the central factor in fair use analysis, explicitly saying that whether or not a use is transformative is in decent part dependent on the degree to which it replaces the original. So if you're writing a generative AI tool that will generate stock photos that it generated by scraping stock photo databases... I mean, the fair use analysis need consist of nothing more than that sentence to conclude that the use is totally not fair; none of the factors weigh in favor it.

Re: Judge dismisses DMCA copyright claim in GitHub Copilot suit

#176
post #164

Earlier quoted context omitted.

What about the copyrights purpose of furthering the arts and sciences?

That ship sailed long ago. While copyright can and is used at times to protect the "little guy", the law is written as it is in order to protect and further corporate interests. The current manifestation of copyright is about rent-seeking, not promoting innovation and creativity. That it may also do so is entirely coincidental.

Also, if it wasn't about rent-seeking and preventing access to works, copyright wouldn't have to last for decades, many multiples of a work's useful commercial life. The fact that it does last this long shows that it's not about promoting innovation and creativity.

Re: Judge dismisses DMCA copyright claim in GitHub Copilot suit

#177

Earlier quoted context omitted.

> How is it any different when a machine does the same thing? Because intent matters in the law. If you intended to reproduce copyrighted code verbatim but tried to hide your activity with a few tweaks, that's a very different thing from using a tool which occasionally reproduces copyrighted code by accident but clearly was not designed for that purpose, and much more often than not outputs transformative works.

> clearly was not designed for that purpose, I'm not aware of evidence that support that claim. If I ask ChatGPT "Give me a recipe for squirrel lemon stew" and it so happens that one person did write a recipe for that exact thing on the Internet, then I would expect that the most accurate, truthful response would be that exact recipe. Anything else would essentially be hallucination.

i think you are misconceiving then how LLMs work / what they are

You can certainly try to hit a nail with a screw driver, but that doesn't make the screw driver a hammer.

Re: Judge dismisses DMCA copyright claim in GitHub Copilot suit

#178
post #152

Earlier quoted context omitted.

Copyright has utterly failed to serve that purpose for a long time, and has been actively counterproductive. But if you want to argue that copyright is counterproductive, I completely agree. That's an argument for reducing or eliminating it across the board, fairly, for everyone; it's not an argument for giving a free pass to AI training while still enforcing it on everyone else .

Without copyright, entire industries would've been dead a long time ago, including many movies, games, books, tv, music, etc. Just because their lobbies tend to push the boundary of copyright into the absurd doesn't mean these industries aren't worth saving. There should be actually respectful lawmakers who seek for a balance of public and commercial interests.

> Without copyright, entire industries would've been dead a long time ago, including many movies, games, books, tv, music, etc.

Citation needed. There are many ways to make money from producing content other than restricting how copies of it can be distributed. The owner should be able to choose copyright as a means of control, but that doesn't mean nobody would create any content at all without copyright as a means of control.

Re: Judge dismisses DMCA copyright claim in GitHub Copilot suit

#179

Earlier quoted context omitted.

> Sometimes the little guy is actually wrong. He is, sometimes. Also sometimes, the moon passes exactly between the sun and Earth, a new star appears in the sky, the magnetic field of our planet reverses, a proton decays (jury is still out on that one, actually). Etc. Tools like Copilot are plagiarism machines. We know the data they're being trained on, and a conclusion of "that's plagiarism" is not - or anyway shoul…

> big guys gang up on little guys all the time And obnoxious individuals gum up enterprises. It's lazy to the point of dismissal to conclude based on bigness.

You can't predict right or wrong based on bigness, but you can very often predict who will win.

EDIT: And by "win" I mean not who the judge will side with, but who will end up chugging along fine financially and who will end up broke.

Re: Judge dismisses DMCA copyright claim in GitHub Copilot suit

#180
post #167
post #136

Earlier quoted context omitted.

I don't think that's something you can take away from the little-guy big-guy narrative. Class actions are funded by courts awarding lawyers huge payouts if they win, not directly by the plaintiffs. There should be plenty of resources on both sides of this fight.

You are sorely underestimating the legal resources available to one of the most powerful companies on earth

I don't believe I am. To flush out my statement more fully there are diminishing returns on investing more money into a lawsuit, and both sides in a class action with this much money at stake should be sufficiently funded to be far beyond the point of diminishing returns.

I'm not claiming Microsoft doesn't have tons of resources, I'm claiming that the plaintiffs attorneys should be sufficiently funded that the difference in outcomes is negligible.

Post reply on HN