> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…
Let's take a second to remember that copyright is the reason ~every child doesn't have access to ~every book ever written. While it might be too disruptive to eliminate copyright overnight, we should remember that our world will be much better and improve much faster to the extent we can reduce copyright's impact. And we should cheer it on when it happens. A majority of the world's population in 2023 has a smartphone…
Sarah Silverman is suing OpenAI and Meta for copyright infringement
491–500 of 599 posts
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#492Earlier quoted context omitted.
Not when OpenAI publicly declared they trained on pirated works. I can’t imagine “we can’t tell if this is the result of the illegal thing we did or not” is going to stand up very well, nor does it bode well for any refutation of the plaintiff’s depiction of their intent. Part of fair use consideration is commercial impact and when you steal a bunch of books to train your AI model, it’s hard to refute that the impact…
… it’s hard to refute that the impact is not negative or that you didn’t intend commercial harm. So this is another thing I don’t understand. Is the claim that fewer people will buy Silverman’s book because ChatGPT is able to provide a summary? If so, call me skeptical.
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#493Earlier quoted context omitted.
That's not was he's saying at all. He's saying you can train an AI on copyrighted material just like people can learn from copyrighted material. If you acquire the material illegally that a separate issue that training AI doesn't give you any protection against.
But acquiring the material illegally is the thrust of the issue, and the point of the thread here: the notion that large companies can get away with piracy if they just execute it on a large enough scale. Copyright infringment for thee but not for me.
'Acquiring' is more difficult to pursue legally. It's easier to go after distribution. In this case, Meta or OpenAI did not distribute anything because they are not chumps. They can go after whoever posted the dataset containing books. Not sure if that is eleuther or just some random person on the internet. In either case, the strategy of going after the rich companies won't work.
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#494> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…
Let's take a second to remember that copyright is the reason ~every child doesn't have access to ~every book ever written. While it might be too disruptive to eliminate copyright overnight, we should remember that our world will be much better and improve much faster to the extent we can reduce copyright's impact. And we should cheer it on when it happens. A majority of the world's population in 2023 has a smartphone…
"Japan's incredibly strong economy is responsible for the manufacture of Datsun cars, boombox stereos, and touch-tone phones..."
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#495Earlier quoted context omitted.
I've followed the issue in the US since the early 2000s as an activist and policy expert. I'm not familiar with the state of play outside the US, but the US is one of the stricter jurisdictions in this regard, for reasons that have mostly to do with sophisticated corruption. I'm responding to the "for me not for thee" and the top comment about there being an inconsistency between the treatment of large companies and…
Are you forgetting all the DMCA lawsuits slapping individuals who downloaded MP3s with tens of thousands of dollars? These were not corporations, these were teenagers still living with their parents who pulled music files off the likes of Napster. The DMCA does allow harassment by copyright holders to individuals suspected of infringement. It's just that most people like authors wouldn't blow their legal budget suing…
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#496Earlier quoted context omitted.
I think that would wholly destroy the ability of writers to actually make a living writing novels. The fact that the living isn't great now does not justify this.
People wrote great works before copywrite. People write for reasons other than money from the sales of the book. This accounts for most authors, who aren't famous enough to negotiate a great deal with a publisher. And we don't need to abolish copyright outright. Just require that it is continually published at a steady or decreasing price, or it becomes public domain. And put works in the public domain a little soone…
Yes, wealthy aristocrats wrote whatever they wanted and less wealthy authors wrote what they got paid to write by their wealthy aristocrat patrons.
Copyright and the publishing industry changed that to make it possible to live by writing for ordinary people.
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#497Are we all reading the same complaint? They say: > in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private tracker.” Does that stack up? The Meta Paper -…
Yes, and no.
Its pretty much a slam dunk case that, insofar as the initial training required causing a copy of the corpus defined by the tracker to be made as part of the process, it involved an act violating copyright.
Whether that entitles Silverman to any remedy beyond compensation for (maybe treble damages for) the equivalent of the purchase price of the book depends on... well, basically the same issues of how copyright relates to model training (and an additional argument about whether the illicit status of the material before the training modifies that).
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#498I wish these cases were stronger than "it summarizes my books".
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#499Earlier quoted context omitted.
Let's take a second to remember that copyright is the reason ~every child doesn't have access to ~every book ever written. While it might be too disruptive to eliminate copyright overnight, we should remember that our world will be much better and improve much faster to the extent we can reduce copyright's impact. And we should cheer it on when it happens. A majority of the world's population in 2023 has a smartphone…
There are already many more public domain books than people are inclined to read: https://www.gutenberg.org/ https://librivox.org/ many of which form the basis for an education: https://news.ycombinator.com/item?id=34630153 And most of which, when in copyright, paid their authors quite handsomely in terms of royalties. If you believe that books should exist without copyright, then one has to ask --- how many books ha…
“To promote the progress of science and useful arts, by securing for limited times to authors and inventors the exclusive right to their respective writings and discoveries.”
If the current copyright scheme does not promote the progress of science and useful arts, it is not performing as intended. All these copyright extensions do little to promote progress!
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#500This is actually quite interesting, as it's drawing a distinction between training material that can be accessed by anybody with a web browser (like anybody's blog), vs. training material that was "illegally-acquired... available in bulk via torrent systems." I don't think there's any reason why this would be a relevant legal distinction in terms of distributing an LLM -- blog authors weren't giving consent either. H…
Eventually, I imagine a new licensing concept will emerge, similar to the idea of music synchronization rights -- maybe call it "training rights." It won't matter whether the text was purchased or pirated -- just like it doesn't matter now if an audio track was purchased or pirated, when it's mixed into in a movie soundtrack. Talent agencies will negotiate training rights fees in bulk for popular content creators, wh…