Live data from Hacker News

Sarah Silverman is suing OpenAI and Meta for copyright infringement

theverge.com

441–450 of 599 posts

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#441

Earlier quoted context omitted.

The benefit of the existing scheme is that new works get created, and some of them are even copy-edited and published. How many texts are created which are explicitly placed in the public domain and from which the authors have made a conscious decision not to profit thereby?

Books and libraries have existed for thousands of years. It’s copyright that is the young intruder. Most people who write non-fiction books do it because they want to contribute to human knowledge and be recognized as an expert in a particular field, not because they think that writing a differential geometry textbook is their path to riches. With the internet, more and more text books are made freely available by th…

And books and libraries existed under a system of patronage and royal imprimatur which severely restricted access, even beyond the expenses of copying a book by hand or physically printing a copy using metal type.

The problems with authors making texts available directly are:

- no gate-keeping, so it's hard to find what is worth reading and what isn't

- no proofreading --- it kills me that errors in books are so casually accepted these days

- few authors have the skills to draw illustrations so as to have a meaningful and clear presentation

I've worked with raw author manuscripts --- in most instances they're not something anyone would choose to read given any other option.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#442
post #392

Earlier quoted context omitted.

The benefit of the existing scheme is that new works get created, and some of them are even copy-edited and published. How many texts are created which are explicitly placed in the public domain and from which the authors have made a conscious decision not to profit thereby?

> The benefit of the existing scheme is that new works get created, and some of them are even copy-edited and published. How do you know that this benefit wouldn't exist in other schemes? Look at permissive open source software which is essentially public domain + shield from liability. No copyright does not mean no compensation. It just means different compensation that doesn't deprave other people of their right to…

And for folks who want to create such works, they are welcome to --- as noted else thread, I've put a fair bit of effort into permissively licensed texts --- but I haven't seen a workable method put forth, nor a rational justification for destroying the value of works recently created under the current system.

I still don't see how a person having copyright over the work which they have created and the ability to license it to their best profit is a harm.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#443
post #28

Isn’t it much more likely that there are a lot of book reviews and summaries in its training set from which it can synthesize its own?

Except they have a documented paper trail showing illegal book repos were used in training

Link? Also have to prove they used all books and did 0 curation to remove copyrighted material

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#444

Isn’t it much more likely that there are a lot of book reviews and summaries in its training set from which it can synthesize its own?

I think it's actually much more likely that they just dumped a bunch of book PDFs in the training folder and let it go to work. I seriously doubt any of these AI companies are being even the least bit careful about the data they're lapping up for training

Training data quality is hugely important so I very much doubt they don’t curate the text they use.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#445

Earlier quoted context omitted.

That's not was he's saying at all. He's saying you can train an AI on copyrighted material just like people can learn from copyrighted material. If you acquire the material illegally that a separate issue that training AI doesn't give you any protection against.

Can I memorize copyrighted material and recite it on Youtube? What if I do so but imperfectly? Where do you draw the line? If it's infringement for a human to do that why is it not for a LLM?

Look at how many people get blocked or demonetized for covers of existing songs which may not even be that close to the original. You can't even play a few seconds of the original in many cases, even for fair use and criticism/discussion. This is already in place.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#446

Earlier quoted context omitted.

Not really. The "special" quality attached to humans is only in creating copyright -- it has nothing to do with fair use arguments around derivative works. "Machines and algorithms are not legally recognized as being able to author original non-derivative works" Neither are monkeys. This doesn't mean a monkey's painting is any more or less derivative, or any more or less subject to a copyright claim. It only means th…

Monkeys aren't algorithms nor computers, so that doesn't seem very relevant. Let's look at a totally different analogy: compression algorithms. If I take a digital artist's work which they publish as a png or psd file, and I use some algorithm to convert it to a jpg file, well, I definitely transformed the work in terms of bytes. It's a smaller file, I threw out a lot of data, you can't get the original back. Yet, th…

" Monkeys aren't algorithms nor computers, so that doesn't seem very relevant."

Works produced by monkeys, like works produced by computers, cannot be copywriten. They are functionally identical in this regard. That is why it's relevant.

This is how the current law works. The analogies will break down with AGI and new law will need to be created. Where LLMs fit into this process is an open question.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#447

Earlier quoted context omitted.

How about you go figure out this new economic model, and come back when it's ready. Until then, the existing model will persist, thank you

Oh boy, do I have good news for you! There are already many writers making thousands of dollars a month by publishing free serialized web novels, via Patreon. Some are using their own websites, but most are on Royal Road (or scribblehub, webnovel, wattpad, AO3). A random example from Royal Road[1], the author makes $12065/month. Mind you, the text is not gated, it's free to read, the patreon only offers early access.…

Your "random" example is currently the #2 ongoing fiction on there, and was #1 until a couple days ago when the guy who wrote Mother of Learning released his new fiction.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#448
post #281

Earlier quoted context omitted.

It seems pretty easy to prove that, since they admitted it in public. Read the article. This isn't about the question of LLMs being copyright infringement, this is about Meta and OpenAI admitting that they had pirated copies of those books.

> > But I really don't see how you could prove OpenAI did that > It seems pretty easy to prove that, since they admitted it in public. Can you highlight/link to where OpenAI have admitted this? As far as I'm aware, OpenAI are still secretive about their training datasets.

It's in the LLaMa.cpp original research paper. Tis mentioned in the brief. The research paper basically stated it was trained on Bibliotik, and other Internet "shadow library" corpuses.

See the reference to the Gao et al, in the linked paper from the article.

Paper linked in article: https://arxiv.org/pdf/2302.13971.pdf

The LLaMA paper references a paper utilizing a data source compiled by EleutherAI otherwise known as ThePile. URL from the bibliography for that paper points yonder: https://zenodo.org/record/7413426

This act of summarization done in a lovingly amateur nature at no cost to you, by domeone who despises copyright in all it's forms, but despises profit oriented self-referential inconsistency by large enterprises even more so.

It's kind of funny, because the more I look into it, the more companies building offerings around stuff like CoPilot, LLaMA, ChatGPT, etc... are pulling something not altogether dissimilar to a Sovereign Citizen trying to worm their way out of a speeding ticket.

They want the benefits of the ML model being trained on no strings attached data corpora, while shirking the obligations that come from operating as a corporate entity in the United States.

Twould be interesting to see if Silverman's legal team can catch Big Tech with their pants down, in a court of law, by pointing this out.

It's really weird. I'm completely split and inable to live with a decision either way in this case due to knock on consequences.

I don't want the likes of OpenAI/Meta/Microsoft/Github getting off without reapong the painful fruits of their own IP related crusades on the sanctity of copyright.

On the other hand, as much of a karmic stiffy as that former outcome gives me, I really want copyright such as it is to die, because computing in general will never be as free as it should be until it does.

This is one of those rare times in life where I'd love to get paid to get locked in a room with judges/legislators to really get it all figured out, because I really don't think that leaving this up to common law jurisprudence is actually the best way to go since the network of knock-on effects are so dramatic in scale.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#449
post #387

Earlier quoted context omitted.

Intellectual property monopolies limit access and utilization of ideas and tools to broader society in exchange for privileging a small group worshipped as "the creators".

That sounds like a good argument, but I'm not seeing how protecting a comedian's IP translates to limiting access to ideas and tools... that benefit broader society.

The same thing that protects that comedian's IP, is the same thing that makes the source designs for lithographic masks for semiconductor fabrication some of the most sensitive IP in the world!

Who it is making the claim doesn't matter half as much as what the knock on consequences to jurisprudence at large.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#450
post #425

A more convincing exhibit would have been convincing ChatGPT to output some of the text verbatim, instead of a summary. Here's what I got when I tried: I'm sorry for the inconvenience, but as of my knowledge cutoff in September 2021, I don't have access to specific external databases, books, or the ability to pull in new information after that date. This means that I can't provide a verbatim quote from Sarah Silverma…

GPT is a lossy jpeg of the whole Internet. It’s not possible to extract verbatim text from it, due to how neural networks work. How do you think they would fit exabytes of text data into a gigabyte-sized neural network? That’s right, it’s lossy.

>It’s not possible to extract verbatim text from it

I didn't ask for the whole book, I asked for the first paragraph. It absolutely is possible to get verbatim text from chatgpt.

Post reply on HN