Live data from Hacker News

Sarah Silverman is suing OpenAI and Meta for copyright infringement

theverge.com

581–590 of 599 posts

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#581

Earlier quoted context omitted.

> It is the equivalent of making a 3D map of a museum and getting sued by one artist of one painting in the museum. If the painting is copyrighted (rather than public domain, as many pieces in museums are), and the map includes an image of that painting, I would expect that to be prohibited. I would prefer the world in which copyright doesn't exist, but while it exists, it should apply to everyone equally.

Sure, but your analogy fails there since it implies that the map contains the entire original work. AI models do not contain the entire original works of their training data. If they did they would be the most efficient storage methods ever devised. All those terabytes of training data can obviously not be squished 1 for 1 into the model that is only a couple of gigabytes at the higher end.

The AI models can, in many cases, reproduce large portions of the training data verbatim. (Some models have controls on top preventing this, but the underlying model has the data.)

And even if they can't reproduce the entire work, they're still derivative works of the training data.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#582

Earlier quoted context omitted.

> You can't find a copy of any specific work anywhere in the weights. You won't find a copy of any specific work anywhere in the compressed form of a file, either, but when you decompress it you find the complete work. And many large AIs can recite, verbatim or near-verbatim, many complete works. Yes, they might get a word wrong, but that doesn't nullify the point that they're trained on the entire work and to a firs…

> You won't find a copy of any specific work anywhere in the compressed form of a file, either Sure you will. It's right there, in PNG encoding or what have you. With nothing more than the compressed file and general purpose tools you can reliably put it on your screen. > And many large AIs can recite, verbatim or near-verbatim, many complete works. This is not the common case and it's not even clear that the reason…

> Those are different people than the holder of the copyright on an arbitrary piece of the training data.

No, I'm saying many AI models use copyrighted works by independent artists to directly compete with those artists (in addition to other artists). And AI models use copyrighted works by software developers to compete with those software developers (in addition to other developers). Fair use determinations do care if you're competing with the work you copied.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#583

> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…

Let's take a second to remember that copyright is the reason ~every child doesn't have access to ~every book ever written. While it might be too disruptive to eliminate copyright overnight, we should remember that our world will be much better and improve much faster to the extent we can reduce copyright's impact. And we should cheer it on when it happens. A majority of the world's population in 2023 has a smartphone…

But it's not children downloading the books is it? It's a company backed by billionaires, so why cheer?

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#584
>did not consent to the use of their copyrighted books as training material

This is an interesting twist on copyright. Boiled down, it more or less summarizes how well copyright is working in 2023:

     X:  Not fair! Your machine looked at my words and used them to make new words!
     Y:  You published it! That means you're sharing!
     X:  You can't have my words because I made these words and I say so!
     Y:  You have to share when you publish! If you won't share, I'm telling on you!
     X:  No, I'm telling on you first!

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#585
post #575

Earlier quoted context omitted.

Replace it with any subject. The point stands - it isn't at all clear how the courts will treat this. Take your example: I'm a self-taught artist, and I learned everything I know about art by studying cartoons made by Disney. Maybe I paid for these cartoons, maybe I didn't. I then make a website where I draw my own cartoons, which, since I've never seen any other art, look a lot like Disney's. Unless I'm straight-up…

> I then make a website where I draw my own cartoons, which, since I've never seen any other art, look a lot like Disney's. Unless I'm straight-up copying their characters, they would have no claim against me. The law pertains ultimately to the actions of humans. We don't allow non-human animals or machines access to legal system. Even in the specious only-Disney-inspired-artist scenario presented (courts don't use u…

Can you point me to some case law that supports your claim?

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#586

Earlier quoted context omitted.

> Let's take a second to remember This is emotionally manipulative speech that provides no value to HN and only serves the purpose of bypassing peoples' logical reasoning circuits. > ~every child doesn't have access to ~every book ever written More manipulation - "think of the children!" Copyright exists because people who produce content with low distribution costs (e.g. books) need some protection for their work be…

> and it's deeply morally wrong (theft-adjacent) to take someone else's work without compensating them on their terms. But I'm willing to bet that you don't believe this consistently across domains, and the domains in which you do believe it have been selected rather arbitrarily, not by you but rather by industry lobbying pressure. Copyright doesn't exist for mathematics, jokes, fashion designs, architectural styles,…

You’re cherry picking a single line and deflecting with whataboutism. This is not benefitting HN.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#587

> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…

Copyright doesn't apply when it comes to fair use, and one of the major factors of fair use is if your use deprived the copyright owner of sales. Good luck arguing that any of the books in question lost sales because an AI was trained on them.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#588

Earlier quoted context omitted.

Sure! Yes! I agree! 100 years is way too long. 20 years is much more reasonable. But the comment that I was responding to (and many others in this thread) are advocating for the complete removal of copyright, and that's what I'm responding to.

100 years is arguably “unconstitutional.” Constitution days “for a limited time”. Death of author + 99 is effectively unlimited to the perspective of typical human

Corporations have personhood in the USA, no?

For the life of a human, yes, for the life of a corporation, it's long but not too long. Specially for corporations such as Disney which are sure to last for quite a while.

Not saying this is a justificable position. I'm actually in favour of drastically reducing copyright (I believe 7 to 15 years might be a sweetspot). But a lot of laws are not made for regular human people any longer.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#589

Earlier quoted context omitted.

> Nothing was stolen- just copied. This typical semantic-pedantry line from piracy apologists misses the point - piracy is theft-adjacent even if you get to pick your use of "theft". Incidentally, my definition of "theft", and that of most content creators, includes the act of consuming something without compensating the creator on their terms - which includes piracy.

A fundamental property of theft is that the action deprives someone of something they had. Piracy is not theft-adjacent. It has nothing do to with theft. It's a completely different concept.

No, a fundamental property of theft is that it is the taking of that which does not belong to you - which clearly includes piracy.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#590
post #544

Earlier quoted context omitted.

> Nothing was stolen- just copied. This typical semantic-pedantry line from piracy apologists misses the point - piracy is theft-adjacent even if you get to pick your use of "theft". Incidentally, my definition of "theft", and that of most content creators, includes the act of consuming something without compensating the creator on their terms - which includes piracy.

[flagged]

"Pretty wild" - standard emotional manipulation. Not appropriate on HN.
Post reply on HN