Live data from Hacker News

Sarah Silverman is suing OpenAI and Meta for copyright infringement

theverge.com

501–510 of 599 posts

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#501
post #489

Earlier quoted context omitted.

People wrote great works before copywrite. People write for reasons other than money from the sales of the book. This accounts for most authors, who aren't famous enough to negotiate a great deal with a publisher. And we don't need to abolish copyright outright. Just require that it is continually published at a steady or decreasing price, or it becomes public domain. And put works in the public domain a little soone…

People did write before copyright. Copyright was established to make it more likely. Yes, people write for other reasons (e.g., self-promotion). But that does not account for "most authors" I would actually want to read. I'm not against copyright being different, especially for things that are written as work-for-hire. But that's a fight with Disney. Good luck.

> People did write before copyright. Copyright was established to make it more likely.

No, it was established to make investing in printing presses more profitable, which is why the rights initially attached to printers. Authors as the locus of rights were a later change.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#502
post #307

Earlier quoted context omitted.

Let's take a second to remember that copyright is the reason ~every child doesn't have access to ~every book ever written. While it might be too disruptive to eliminate copyright overnight, we should remember that our world will be much better and improve much faster to the extent we can reduce copyright's impact. And we should cheer it on when it happens. A majority of the world's population in 2023 has a smartphone…

Then most people stop writing books because they can't get paid for their time/effort and ~every child will be stuck with outdated knowledge within a decade.

The existence of Wikipedia seems to be a pretty strong counterargument to this.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#503

Earlier quoted context omitted.

I've followed the issue in the US since the early 2000s as an activist and policy expert. I'm not familiar with the state of play outside the US, but the US is one of the stricter jurisdictions in this regard, for reasons that have mostly to do with sophisticated corruption. I'm responding to the "for me not for thee" and the top comment about there being an inconsistency between the treatment of large companies and…

Are you forgetting all the DMCA lawsuits slapping individuals who downloaded MP3s with tens of thousands of dollars? These were not corporations, these were teenagers still living with their parents who pulled music files off the likes of Napster. The DMCA does allow harassment by copyright holders to individuals suspected of infringement. It's just that most people like authors wouldn't blow their legal budget suing…

My understanding is that the issue there was that they were simultaneously sharing those MP3s, as they were downloading them. Which is to say, I think the person you're responding to is correct about the difference between sharing and downloading, but I'm not a lawyer.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#504

Earlier quoted context omitted.

> If everyone is allowed to steal books Nothing was stolen- just copied.

[flagged]

That's not really the same, you could have a copy of my credit card, but you using it to make purchases would become an issue. Regardless, that quickly steps out of the domain of intellectual property.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#505

Earlier quoted context omitted.

It is the equivalent of making a 3D map of a museum and getting sued by one artist of one painting in the museum. Ant individual work in an AI dataset is nearly worthless - only in aggregate does it have value. If that doesn't count as a "transformative work" I don't know what does.

> It is the equivalent of making a 3D map of a museum and getting sued by one artist of one painting in the museum. If the painting is copyrighted (rather than public domain, as many pieces in museums are), and the map includes an image of that painting, I would expect that to be prohibited. I would prefer the world in which copyright doesn't exist, but while it exists, it should apply to everyone equally.

Sure, but your analogy fails there since it implies that the map contains the entire original work. AI models do not contain the entire original works of their training data. If they did they would be the most efficient storage methods ever devised. All those terabytes of training data can obviously not be squished 1 for 1 into the model that is only a couple of gigabytes at the higher end.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#506
post #396
post #385

Earlier quoted context omitted.

Is possession of a pirated work the same as possession of stolen property, legally?

Absolutely it is. There is two centuries of precedence marking copyrighted works as property. It's literally called intellectual property . The courts have only made clarification that intellectual property doesn't violate physical property theft laws (denying ownership), but instead intellectual property theft laws (denying compensation).

But "possession of stolen property" is a specific criminal charge that doesn't typically have to do with copyright (being more to do with, say, bicycles). I can't find any example of "possession of stolen property" charges being brought against a copyright violator, but it's a hard thing to Google.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#507
post #392

Earlier quoted context omitted.

> The benefit of the existing scheme is that new works get created, and some of them are even copy-edited and published. How do you know that this benefit wouldn't exist in other schemes? Look at permissive open source software which is essentially public domain + shield from liability. No copyright does not mean no compensation. It just means different compensation that doesn't deprave other people of their right to…

> Look at permissive open source software ... No copyright does not mean no compensation From everything I've heard it kinda does. If you're writing something valuable then maybe a company will employ you to keep working on it, and the portfolio can certainly help in interviews (to write other software), but getting non-negligible compensation for the use of the software itself is rare. Even those projects that are w…

> From everything I've heard it kinda does. If you're writing something valuable then maybe a company will employ you to keep working on it, and the portfolio can certainly help in interviews (to write other software), but getting non-negligible compensation for the use of the software itself is rare. Even those projects that are well funded, like the Linux kernel, are done so not out of goodness of heart, but due to companies realizing it's in their rational interest to have a common standard base of sorts.

Successfully creating a permissive open source project expecting it to be magically funded by benevolent parties is just exceedingly rare. Usually the funding starts first, or effort proceeds in lock-step with funding.

There's a bit of a cart-and-horse here though. Non-permissive licenses like AGPL are often not the result of a single author hoping they might be able to negotiate some licensing deal in the future - it is the result of a commercial enterprise trying to be restrictive in the ability for people to use their source without compensating them.

Same with what is normally considered a more open license, than AGPL, the GPL - where MySQL was reported as going after others for using independent database drivers without buying a commercial license, saying use of the MySQL network protocol made the application using the driver a "derivative work" under the GPL.

Linux is a special case because there are a large number of commercial entities which realize contributing to Linux is way cheaper and faster than writing their own kernel and porting user land software to it.

Apache HTTPD, on the other hand, is an example of an application where corporations DID find the motivation to write their own alternative funded with a commercial model, such as NGINX.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#508
post #281

Earlier quoted context omitted.

> > But I really don't see how you could prove OpenAI did that > It seems pretty easy to prove that, since they admitted it in public. Can you highlight/link to where OpenAI have admitted this? As far as I'm aware, OpenAI are still secretive about their training datasets.

It's in the LLaMa.cpp original research paper. Tis mentioned in the brief. The research paper basically stated it was trained on Bibliotik, and other Internet "shadow library" corpuses. See the reference to the Gao et al, in the linked paper from the article. Paper linked in article: https://arxiv.org/pdf/2302.13971.pdf The LLaMA paper references a paper utilizing a data source compiled by EleutherAI otherwise known…

> It's in the LLaMa.cpp original research paper

To my understanding:

* Wowfunhappy said "I really don't see how you could prove OpenAI did that", verve_rat replied "they admitted it in public", I asked "where OpenAI have admitted this" and noted "OpenAI are still secretive about their training datasets" - specifically about the OpenAI claim

* LLaMA(.cpp) is (an unofficial implementation of) Facebook's leaked model

On balance of probabilities I'd guess that OpenAI did train on material not legally acquired, but as far as I'm aware they've never actually admitted to what's in their dataset as is being claimed.

> It's really weird. I'm completely split and inable to live with a decision either way in this case due to knock on consequences.

I think strengthening of IP law risks hindering the field (of which the majority is uncontroversially positive but too boring for press attention, like defect detection, language translation, spam/DDoS filtering, agriculture/weather/logistics modelling, etc.) while still ending up hurting individuals and FOSS/academic research more than those with large data moats (Microsoft with Github repos, Google with Youtube videos, Adobe and Getty with stock images, etc.)

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#509

Earlier quoted context omitted.

It would look largely identical to ours, I think. It's pretty trivial to get access to many, if not most, e-books. Any public-domain work is available on Project Gutenberg [0]. Copyrighted works can be accessed for free using tools of various legality: Libby [1] is likely sponsored by your local library and gives free access to e-books and audiobooks. Library Genesis [2] has a questionable legal status but has a huge…

Public domain would be fine, if we had the original copyright term instead of "life of the author plus seventy years." Libby is an interesting option, though I'm curious how many kids in disadvantaged countries would actually have access to it. Regarding Libgen, I'm not convinced it makes the case for modern copyright to say it's fine, because people can just violate copyright.

Exactly. The original copyright term was 14 years. Now Disney and the Intellectual Monopoly cartel have extended it 120 years, and convinced an unthinking population that anything less is "stealing."

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#510

> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…

I am not sure this is exactly correct. If you download a book that might be copyright infringement. But not if I download a word. How much do they need to download at a time before it becomes infringement? And if the material is never used or displayed to a human is it still infringment (if so, Google is awaiting a huge lawsuit)? Alternatively if I, a human, read a book it is copied into my memory. Is that infringment? What if I quote it? How much can I quote at what frequency before I'm infringing? If I write something similar to the book but in my own words, is that infringment? How similar does it need to be? What about derivative works and fair use?

Copyright is a horrible mess of individual judgements and opinions. Written material especially. And the same applies to AI. So now we will get a judgement which is a tech-illiterate judges best guess at the intention of a law written to deal with printing presses not AIs and no room for nuances.

Post reply on HN