Live data from Hacker News

Sarah Silverman is suing OpenAI and Meta for copyright infringement

theverge.com

521–530 of 599 posts

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#521

Earlier quoted context omitted.

> If everyone is allowed to steal books Nothing was stolen- just copied.

As a book author, I can say that, "Yes, something was stolen. My opportunity to earn a living taking care of readers." Now you may believe the incredibly self-serving baloney from big companies like Google. You may want to pretend that infringement isn't theft. To you, I hope that some homeless kid breaks into your home, starts squatting, and says that, "Hey, this isn't theft. Nothing has been destroyed."

[flagged]

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#522

Earlier quoted context omitted.

There are already many more public domain books than people are inclined to read: https://www.gutenberg.org/ https://librivox.org/ many of which form the basis for an education: https://news.ycombinator.com/item?id=34630153 And most of which, when in copyright, paid their authors quite handsomely in terms of royalties. If you believe that books should exist without copyright, then one has to ask --- how many books ha…

The good public domain books are typically outdated in both content and language. This makes it hard for those with less resources to stay competitive and makes the task of understanding unnecessarily hard. The link you posted that you say used public works to form a “basis for an education” uses Aristotle as an author for example, and seem to be taught in the context of an instructor-led class (where an expert can d…

The problem with your reasoning is similar to coming with the highest possible number. When you figure out a number, someone will say "that plus one!".

Whatever standard of free service you establish, someone with resources can surpass, but needs to be motivated to do it by the perspective of a return of the resources spent.

So either hinder development by forbidding the commercial stuff altogether, or tax all people and then finance authors from public money, but hopefully I don't have to explain how wherever this model is tested, the society degrades towards famine.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#523

Earlier quoted context omitted.

While I’ma proponent of free information and loosening copyright, allowing billion dollar companies to package up the sum of human creation and resell statistical models that mimic the content and style of everyone… is a bit far. Fair use is for humans.

Yeah, but hypothetically should open source projects be offered special protections? I feel like they should, and with certain caveats where, say, a company like Meta is allowed to claim fair use if and only if they free up the entire ecosystem as they deploy it. But yeah, having private ostensibly profitable models based on other people's work without giving them free access to it is not fair. Give some get some.

Maybe… maybe restrictions freed up for resources in the public domain… that is no patent, copyright, license or ownership rights of any kind. As in you could distribute the thing unmolested but that’s it.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#524
post #392

Earlier quoted context omitted.

> The benefit of the existing scheme is that new works get created, and some of them are even copy-edited and published. How do you know that this benefit wouldn't exist in other schemes? Look at permissive open source software which is essentially public domain + shield from liability. No copyright does not mean no compensation. It just means different compensation that doesn't deprave other people of their right to…

> Look at permissive open source software ... No copyright does not mean no compensation From everything I've heard it kinda does. If you're writing something valuable then maybe a company will employ you to keep working on it, and the portfolio can certainly help in interviews (to write other software), but getting non-negligible compensation for the use of the software itself is rare. Even those projects that are w…

There's a huge vendor community around Kubernetes, which is open source and permissively licensed to boot. If you write something complex that basically works you'll be able to develop defensible IP around management.

p.s., I'm sure many of those vendors are not making money yet but they all aim to.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#525

> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…

Let's take a second to remember that copyright is the reason ~every child doesn't have access to ~every book ever written. While it might be too disruptive to eliminate copyright overnight, we should remember that our world will be much better and improve much faster to the extent we can reduce copyright's impact. And we should cheer it on when it happens. A majority of the world's population in 2023 has a smartphone…

> Let's take a second to remember

This is emotionally manipulative speech that provides no value to HN and only serves the purpose of bypassing peoples' logical reasoning circuits.

> ~every child doesn't have access to ~every book ever written

More manipulation - "think of the children!"

Copyright exists because people who produce content with low distribution costs (e.g. books) need some protection for their work being taken without compensation.

Fundamentally, you are never entitled to someone else's work.

There's already tens (hundreds?) of thousands of books in the public domain, and tens of thousands more under Creative Commons licenses (where the author explicitly released their work for free distribution). There's lectures on YouTube and MIT OpenCourseWare. There's full K-12 textbooks on OpenStax and Wikibooks. There's Wikipedia, Stack Exchange, the Internet Archive, and millions of small blogs and websites hosting content that is completely free.

There is no need for "a majority of the world had access to every book ever digitized" - and it's deeply morally wrong (theft-adjacent) to take someone else's work without compensating them on their terms.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#526

Earlier quoted context omitted.

I’d more ask what the cost vs benefits are of keeping the existing scheme, it’s not free to run all these DRM services, prosecute offenders etc… Not to say I support no copyright…

The benefit of the existing scheme is that new works get created, and some of them are even copy-edited and published. How many texts are created which are explicitly placed in the public domain and from which the authors have made a conscious decision not to profit thereby?

This is a logical error.

The benefit of automobiles is that people move across vast distances.

Wrong. People used to move across vast distances before, using horses. Yes, automobiles are better and now traveling is easier. But we have no ability to figure out how many people would travel on horses if there wasn't a better alternative. It's even possible, that banning cars could eventually lead to an even better method of transportation (escaping a local minimum kind of a thing).

Authors who write books have tools to protect their interests, so they do so. Without these tools perhaps there would be less authors. Or maybe there would be more authors: I think Windows and Photoshop are so popular because they were pirated a lot; returning to the context of books, less and less people read them, but maybe if books (attractive books, not old books in public domain) were free, then the trend would reverse, books would popularize, people would start enjoying deeper entertainment, get smarter, transform the society for the better, and support authors on e.g. Patreon… Or maybe not, I'm just mentioning some nuance.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#527

Earlier quoted context omitted.

Second-order effects matter, though: If everyone is allowed to steal books, what's the incentive for experts to write new ones, and for the publishers to reward them for it? Btw, not a fan of "but what about the kids" rhetoric: https://en.wikipedia.org/wiki/Think_of_the_children

> If everyone is allowed to steal books Nothing was stolen- just copied.

> Nothing was stolen- just copied.

This typical semantic-pedantry line from piracy apologists misses the point - piracy is theft-adjacent even if you get to pick your use of "theft". Incidentally, my definition of "theft", and that of most content creators, includes the act of consuming something without compensating the creator on their terms - which includes piracy.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#528

Earlier quoted context omitted.

There are already many more public domain books than people are inclined to read: https://www.gutenberg.org/ https://librivox.org/ many of which form the basis for an education: https://news.ycombinator.com/item?id=34630153 And most of which, when in copyright, paid their authors quite handsomely in terms of royalties. If you believe that books should exist without copyright, then one has to ask --- how many books ha…

People keep forgetting the purpose of copyright. It’s easy to find, it’s in the Constitution! “To promote the progress of science and useful arts, by securing for limited times to authors and inventors the exclusive right to their respective writings and discoveries.” If the current copyright scheme does not promote the progress of science and useful arts, it is not performing as intended. All these copyright extensi…

> If the current copyright scheme does not promote the progress of science and useful arts, it is not performing as intended. All these copyright extensions do little to promote progress!

This is a strawman argument - the parent poster that you're replying to is defending copyright as a concept, while the one that they're replying to is attacking it - you're instead making an argument about specific lengths that hasn't come up yet, and nobody is defending.

Very few people (and I am not one of them) think that the "Mickey Mouse curve" style of copyright extension is genuinely useful to anyone except Disney, but holmesworcester is arguing that copyright should be abolished entirely, which contradicts the section of the Constitution that you quoted.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#529

Earlier quoted context omitted.

How about you go figure out this new economic model, and come back when it's ready. Until then, the existing model will persist, thank you

Oh boy, do I have good news for you! There are already many writers making thousands of dollars a month by publishing free serialized web novels, via Patreon. Some are using their own websites, but most are on Royal Road (or scribblehub, webnovel, wattpad, AO3). A random example from Royal Road[1], the author makes $12065/month. Mind you, the text is not gated, it's free to read, the patreon only offers early access.…

This model is clearly not viable for any sort of economy at scale. Far fewer people can make a living wage under patronage models than traditional models because people simply pay less (far less) for things where they're not compelled to than for things where they are.

420 thousand people work in the US film industry alone[1]. I doubt that there's that same number of people making ~living wage across every online patronage site in the United States (excluding advertisement-driven ventures, of course).

[1] https://www.statista.com/statistics/184412/employment-in-us-...

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#530

> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…

Let's take a second to remember that copyright is the reason ~every child doesn't have access to ~every book ever written. While it might be too disruptive to eliminate copyright overnight, we should remember that our world will be much better and improve much faster to the extent we can reduce copyright's impact. And we should cheer it on when it happens. A majority of the world's population in 2023 has a smartphone…

Why don't children have access to all the latest LLMs, including ChatGPT-4, for free?

Why does money exist?

Post reply on HN