Live data from Hacker News

Judge dismisses DMCA copyright claim in GitHub Copilot suit

theregister.com

471–480 of 505 posts

Re: Judge dismisses DMCA copyright claim in GitHub Copilot suit

#471

Earlier quoted context omitted.

> And I seem to recall there are some theoretical lower bounds on even lossy compression. I'm not sure what your math is coming from and it seems trivially wrong. A single black pixel is a very lossy compression of every image on the internet. A picture of the Facebook logo is a slightly-less-lossy compression of every picture on the internet (the Facebook logo shows up on a lot of websites). I would believe that you…

Ok, that's a really low lower bound. I think you'll agree that it would be a bit absurd to threaten legal action against someone for storing a single black pixel. OTOH Someone might be tempted to start a lawsuit if they believe their image is somehow actually stored in a particular data file. For this to be a viable class action lawsuit to pursue, I think you'd have to subscribe to the belief that it's a form of comp…

I think that when you speak in terms of images, for a viable lawsuit, you need to have a form of compression that can recall n (n >= 1) images from compressing m (m >= n) images. Presumably n is very large for LLMs or image models, even though m is orders of magnitude larger. I do not think that your form of compression needs to be able to get all m images back. By forcing m = n in your argument, you are forcing some idea of uniformity of treatment in the compression, which we know is not the case.

The black pixel won't get you sued, but the Facebook logo example I used could get you sued. Specifically by Facebook. There is an image (n = 1) that is substantially similar to the output of your compression algorithm.

That is sort of what Getty's lawsuit alleges. Not that every picture is recallable from an LLM, but that several images that are substantially similar to Getty's images are recallable. The same goes with the NYT's lawsuit and OpenAI.

Re: Judge dismisses DMCA copyright claim in GitHub Copilot suit

#472

Earlier quoted context omitted.

Ok, that's a really low lower bound. I think you'll agree that it would be a bit absurd to threaten legal action against someone for storing a single black pixel. OTOH Someone might be tempted to start a lawsuit if they believe their image is somehow actually stored in a particular data file. For this to be a viable class action lawsuit to pursue, I think you'd have to subscribe to the belief that it's a form of comp…

I think that when you speak in terms of images, for a viable lawsuit, you need to have a form of compression that can recall n (n >= 1) images from compressing m (m >= n) images. Presumably n is very large for LLMs or image models, even though m is orders of magnitude larger. I do not think that your form of compression needs to be able to get all m images back. By forcing m = n in your argument, you are forcing some…

Thank you for talking with me!

I do realize the benefits of the 'compression' model of ML. Sometimes you can even use compression directly, like here: https://arxiv.org/abs/cs/0312044 .

I suppose you're right that you only need a few substantively similar outputs to potentially get sued already. (depending on who's scrutinizing you).

While talking with you, it occurred to me that so far we've ignored the output set o, which is the set of all images output by -say- stable diffusion. n can then be defined as n = m ∩ o .

And we know m is much larger than n, and o is theoretically practically infinite [1] (you can generate as many unique images as you like) , so o >> m >> n . [2]

Really already at this point I think calling SD a compression algorithm might be just a little odd. It doesn't look like the goal is compression at all. Especially when the authors seem to treat n like a bug ('overfit'), and keep trying to shrink it.

That's before looking back at the "compression ratio" and "loss ratio" of this algorithm, so maybe in future I can save myself some maths. It's an interesting approach to the argument I might try more in future. (Thank you for helping me to think in this direction)

* I think in the case of the Getty lawsuit they might have a bit of a point, if the model might have been overfitted on some of their images. Though I wonder if in some cases the model merely added Getty watermarks to novel images. I'm pretty sure that will have had something to do with setting Getty off.

* I am deeply suspicious of the NYT case. There's a large chunk of examples where they used ChatGPT to browse their own website. This makes me wonder if the rest of the examples are only slightly more subtle. IIRC I couldn't replicate them trivially. (YMMV, we can revisit if you're really interested)

[1] However, in practice there appear to be limits to floating point precision.

[2] I'm using >> as "much greater than"

Re: Judge dismisses DMCA copyright claim in GitHub Copilot suit

#473

Earlier quoted context omitted.

We converged on a system that makes copying illegal because that system was invented in an era when the only people who could copy were those with specialized equipment (e.g. printing presses). In that world, those who might do the copying were often larger than those whose works were being copied, and copyright had more potential to be "protective". That system hasn't been updated for a world in which everyone can m…

> the much larger players who are mass-copying works largely by individuals or smaller entities have become effectively exempt from copyright That's not true. I'm a copyright attorney and I spend my day extracting money from the largest players on behalf of individuals.

I was referring to AI training here.

Re: Judge dismisses DMCA copyright claim in GitHub Copilot suit

#474
post #264

All GitHub needs to do to make most happy is offer an opt-out toggle. It still doesn't.

That wouldn't and shouldn't make most people happy. Repository owner != author for all the code - that's kind of the point of open source.

Good point. But on top of fork hierarchy, git commit authors can be used.

This would mean excluding non-github members, and excluding members that opt out.

Re: Judge dismisses DMCA copyright claim in GitHub Copilot suit

#475

Earlier quoted context omitted.

It's an extreme stretch to say that the model weights are a derivative work of the training data given the legal definition of "derivative work".

It's not more a stretch than saying that re-encoding a PNG as a JPEG is a derivative work even though the process is lossy and the resulting bits look nothing alike.

I'm not sure you're being intellectually honest.

You think that a model that's capable of being prodded into producing an infringing output in addition to all the other non-infringing outputs it could produce is no different than a compression algorithm?

Re: Judge dismisses DMCA copyright claim in GitHub Copilot suit

#476

Earlier quoted context omitted.

Yes there is. If I can emulate Super Mario Odyssey on my PC, I don't need to buy a Nintendo Switch. If it wasn't available there, I'd have to buy a Nintendo Switch to play it. That's a lost sale for Nintendo. You could argue that I wasn't going to buy a switch anyway, but then we're getting too into hypotheticals.

This is the same reasoning the music and movie industries use when they go after people downloading music. And contrary to the popular opinion, I think it is wrong: if people want to pay, they will pay. Same for movies: if people would really want to pay for a movie, they would go to a cinema. Or stream it after a week or two. But there are also people who would jump through hoops than pay for music or movies. And th…

Music isn't video games or movies, and is experienced differently, so while there are similarities, it's not the same because they aren't same thing.

Locks keep people honest. Unfortunately, software lockpicks have the unfortunate reality of being as easily distributable as the software itself.

Re: Judge dismisses DMCA copyright claim in GitHub Copilot suit

#479

Earlier quoted context omitted.

> Sometimes the little guy is actually wrong. He is, sometimes. Also sometimes, the moon passes exactly between the sun and Earth, a new star appears in the sky, the magnetic field of our planet reverses, a proton decays (jury is still out on that one, actually). Etc. Tools like Copilot are plagiarism machines. We know the data they're being trained on, and a conclusion of "that's plagiarism" is not - or anyway shoul…

> big guys gang up on little guys all the time And obnoxious individuals gum up enterprises. It's lazy to the point of dismissal to conclude based on bigness.

Won't someone PLEASE think about microsoft???
Post reply on HN