Live data from Hacker News

Sarah Silverman is suing OpenAI and Meta for copyright infringement

theverge.com

561–570 of 599 posts

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#561

Earlier quoted context omitted.

Let's take a second to remember that copyright is the reason ~every child doesn't have access to ~every book ever written. While it might be too disruptive to eliminate copyright overnight, we should remember that our world will be much better and improve much faster to the extent we can reduce copyright's impact. And we should cheer it on when it happens. A majority of the world's population in 2023 has a smartphone…

Second-order effects matter, though: If everyone is allowed to steal books, what's the incentive for experts to write new ones, and for the publishers to reward them for it? Btw, not a fan of "but what about the kids" rhetoric: https://en.wikipedia.org/wiki/Think_of_the_children

Plenty of freely published articles and fanfics online

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#562

Earlier quoted context omitted.

> otherwise you'll soon find we have far fewer great authors, artists, etc. It's been long understood that this idea isn't based on any real evidence. Creators create because they like to create things. Adding money to the mix tends to ruin most creative endeavours. Look up beautification or note how Googles only good search results involve the keyword Reddit.

Theres "creators create because they like to create things" and "creators create things because they want to make a living off of what they create". If artists/authors/musicians/etc aren't going to be paid for what they create, they can't make a living doing it. If they can't make a living doing it, that severely limits their opportunity and time available for creating things since they could only do it as a hobby (u…

Copyright laws are red flag laws for publishers.

It has nothing to do with creators.

People are regularly paying creators directly to create through patreon, super chats, advertising, early access, subscriptions, etc.

The idea that you need copyright to protect you is just not based in reality.

Get rid of copyright, creators will find a way to monetize it if they want to make a living doing it.

It's not society's job to protect your ability to get paid for your hobby. There are no original ideas, just people that write them down. You don't own them, you extracted these ideas from society, the least you can do is give them back.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#563

> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…

> But companies like Google and Facebook get to play by different rules It's simple, copyrighted materials can be used for academic research. That's what they are doing. Trying new AI modes, publishing results, etc. Facebook doesn't make money on LLaMA, they even require permission to use their models for, again, academic research.

But what if those copyrighted materials were illegally gained? That's what the suit alleges.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#564

Earlier quoted context omitted.

In general copyright (in the us) doesn't cover transformational usage. If you can argue that the nature of your use is transformative you might be good.

"The transformative use concept arose from a 1994 decision by the U.S. Supreme Court. In Campbell v. Acuff-Rose Music, the Court focused not only on the small quantity taken from the copyrighted work but also on the transformative nature of the defendant’s use. The case concerned a song by the group 2 Live Crew entitled "Pretty Woman," which, according to an affidavit, was meant to "through comical lyrics, satirize t…

> or uses an insignificant piece of the work within a larger work with a different aim or purpose

I think this is the crux of the issue, and why I don't see a path to courts ruling that training AI is infringement. My bet is on a Fair Use ruling, though my confidence is not high. As a thought experiment, I considered llama 65B: the 4-bit quantized model is 38.5GB. The model itself was trained on 1.4T tokens, each token being ~4 characters (using OpenAIs stats for English here). Thats 5.6T characters, or 5.09TB of training data. The final model, as a porportion of the total size of the data, is 38.5GB/5090GB = .0075 = 0.7%.

I think it's pretty hard to argue that processing the data and throwing more than 99% of it away means they are "unlimited in the purposes that the work can be put to". Indeed, even replicating a single work using such a model would be enormously difficult.

But returning to your statement regarding the amount used and the purpose: AI models are not competing with books for readers. So I would argue training an AI on these works constitutes fair use, given that the final work (the model) uses less than 1% of the original works, and has a different aim and purpose that the original works.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#565

Earlier quoted context omitted.

Copyright is a system meant to ensure creators get compensated proportionally to the value they add to society It is a bad system though. It restrains the society from benefitting from said work unless they meet certain terms (usually, payments). It's often times unrealistic to pay for all you'd like to consume. Especially that access to their work is usually badly quantised (For example, I may want to search Sarah S…

> Copyright is a system meant to ensure creators get compensated proportionally to the value they add to society Compensation to the creators was merely a means to an end. Furthering progress was the goal: > Article I, Section 8, Clause 8: Patent and Copyright Clause of the Constitution. [The Congress shall have power] “To promote the progress of science and useful arts, by securing for limited times to authors and i…

Well doesn't that just provide extreme support for the cause to reduce copy protection to ensure access for all who could benefit?

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#566

Earlier quoted context omitted.

> Copyright is a system meant to ensure creators get compensated proportionally to the value they add to society Compensation to the creators was merely a means to an end. Furthering progress was the goal: > Article I, Section 8, Clause 8: Patent and Copyright Clause of the Constitution. [The Congress shall have power] “To promote the progress of science and useful arts, by securing for limited times to authors and i…

Well doesn't that just provide extreme support for the cause to reduce copy protection to ensure access for all who could benefit?

Yes

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#567

Earlier quoted context omitted.

> Copyright is a system meant to ensure creators get compensated proportionally to the value they add to society Compensation to the creators was merely a means to an end. Furthering progress was the goal: > Article I, Section 8, Clause 8: Patent and Copyright Clause of the Constitution. [The Congress shall have power] “To promote the progress of science and useful arts, by securing for limited times to authors and i…

Well doesn't that just provide extreme support for the cause to reduce copy protection to ensure access for all who could benefit?

Depends on how you define "progress".

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#568

Earlier quoted context omitted.

It's worth challenging the length of copyright. 20 years seems good enough for high tech R&D, probably good for books as well.

Sure! Yes! I agree! 100 years is way too long. 20 years is much more reasonable. But the comment that I was responding to (and many others in this thread) are advocating for the complete removal of copyright, and that's what I'm responding to.

I was not advocating for the immediate complete removal; I acknowledged that this may be too disruptive.

I was simply reminding people what direction we should go in, and what the stakes are. Reducing copyright terms is a great solution, and yes something between 2 and 12 years is probably the right number to aim for as a first step. I agree with the above commenter that 20 is far too long, because 20-year-old material is slipping into irrelevance in many cases.

It's also good to remember that the only reason we have copyright (in the US, where "moral rights" are not a thing) is to stimulate the creation of more work. So we need to think about what configuration of copyright law actually stimulates more work, and we need to be willing to experiment.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#569
post #496

Earlier quoted context omitted.

People wrote great works before copywrite. People write for reasons other than money from the sales of the book. This accounts for most authors, who aren't famous enough to negotiate a great deal with a publisher. And we don't need to abolish copyright outright. Just require that it is continually published at a steady or decreasing price, or it becomes public domain. And put works in the public domain a little soone…

> People wrote great works before copywrite. Yes, wealthy aristocrats wrote whatever they wanted and less wealthy authors wrote what they got paid to write by their wealthy aristocrat patrons. Copyright and the publishing industry changed that to make it possible to live by writing for ordinary people.

And now wealthy aristocrats have been replaced by staggeringly wealthy megacorporations like Disney and WB.

Its disgustingly dishonest to appeal for "poor" authors where casual glance immediately proves that modern IP law exceedingly profits only a few massive corporations and their shareholders.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#570

Earlier quoted context omitted.

Let's take a second to remember that copyright is the reason ~every child doesn't have access to ~every book ever written. While it might be too disruptive to eliminate copyright overnight, we should remember that our world will be much better and improve much faster to the extent we can reduce copyright's impact. And we should cheer it on when it happens. A majority of the world's population in 2023 has a smartphone…

> Let's take a second to remember This is emotionally manipulative speech that provides no value to HN and only serves the purpose of bypassing peoples' logical reasoning circuits. > ~every child doesn't have access to ~every book ever written More manipulation - "think of the children!" Copyright exists because people who produce content with low distribution costs (e.g. books) need some protection for their work be…

> and it's deeply morally wrong (theft-adjacent) to take someone else's work without compensating them on their terms.

But I'm willing to bet that you don't believe this consistently across domains, and the domains in which you do believe it have been selected rather arbitrarily, not by you but rather by industry lobbying pressure.

Copyright doesn't exist for mathematics, jokes, fashion designs, architectural styles, recipes, and many other areas of human work. All of these represent similar creative work to the work done by musicians and writers. But we don't force comedians to license each others' jokes, or sue bars for letting people tell unlicensed jokes in public. And almost all of us wear clothing by uncompensated designers. And of course it would be unfathomably destructive to allow something analogous to copyright for a mathematical idea.

We also set limits on how long heirs can inherit copyright, which we don't do for other kinds of property, and we don't have any moral issues with that.

So it's important to remember that our moral intuitions about work and material products don't really translate to information and that we are truly making all of this up as we go along, under the intense corrupting pressure of a few very sophisticated industry lobbies.

Post reply on HN