Earlier quoted context omitted.
Isn’t the burden of proof on the other side?
Not when OpenAI publicly declared they trained on pirated works. I can’t imagine “we can’t tell if this is the result of the illegal thing we did or not” is going to stand up very well, nor does it bode well for any refutation of the plaintiff’s depiction of their intent. Part of fair use consideration is commercial impact and when you steal a bunch of books to train your AI model, it’s hard to refute that the impact…
Sarah Silverman is suing OpenAI and Meta for copyright infringement
131–140 of 599 posts
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#132Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#133Earlier quoted context omitted.
> Is there any legal basis for saying fair use permits distributing an LLM trained on copyrighted material, but you have to purchase all the content first to do so legally if it's only available for sale? My understanding (disclaimer: IANAL) is that in order to claim fair use, you have to be legally in possession of the work. If the work is only legally available for sale, then you must have legally purchased a copy,…
> you must have legally purchased a copy, or been given it by someone who did so (for example, if you received it as a gift). I am also NAL, but can I imagine it goes further than that. Just purchasing a copy doesn't let you create and sell (directly as content or indirectly via a service like a chatbot) derivative works that are substantially similar in style and voice to the original work. For example, an LLM 's re…
You're essentially banning satire here, though. There's plenty of folks making a living as cover bands or impersonators. I'm not sure what the answer is, but it's definitely not outright outlawing imitation.
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#134> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…
I for one am quite happy that AI folks are basically treating copyright as not existing. I strongly hope that the courts find that LLM weights, and the datasets are "fair-use" or whatever other silly legal justification. Aaron Swartz was a saint.
Fair use is for humans.
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#135> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…
I for one am quite happy that AI folks are basically treating copyright as not existing. I strongly hope that the courts find that LLM weights, and the datasets are "fair-use" or whatever other silly legal justification. Aaron Swartz was a saint.
> LLM weights, and the datasets are "fair-use" or whatever other silly legal justification.
Would just be a carve out for the wealthy. If these laws don't mean anything, everyone who got harassed, threatened, extorted, fined, arrested, tried, or jailed for internet piracy are owed reparations. Let Meta pay them.
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#136> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…
I for one am quite happy that AI folks are basically treating copyright as not existing. I strongly hope that the courts find that LLM weights, and the datasets are "fair-use" or whatever other silly legal justification. Aaron Swartz was a saint.
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#137> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…
I for one am quite happy that AI folks are basically treating copyright as not existing. I strongly hope that the courts find that LLM weights, and the datasets are "fair-use" or whatever other silly legal justification. Aaron Swartz was a saint.
I would be very happy if either a court or lawmakers decided that copyright itself was unconscionable. That isn't what's going to happen, though. And I think it's incredibly unacceptable if a court or lawmakers instead decide that AI training in particular gets a special exception to violate other people's copyrights on a massive scale when nobody else gets to do that.
As far as a fair use argument in particular, fair use in the US is a fourfold test:
> the purpose and character of the use, including whether such use is of a commercial nature or is for nonprofit educational purposes;
The purpose and character is absolutely heavily commercial and makes a great deal of money for the companies building the AIs. A primary use is to create other works of commercial value competing with the original works.
> the nature of the copyrighted work;
There's nothing about the works used for AI training that makes them any less entitled to copyright protections or more permissive of fair use than anything else. They're not unpublished, they're not merely collections of facts or ideas, etc.
> the amount and substantiality of the portion used in relation to the copyrighted work as a whole; and
AI training uses entire works, not excerpts.
> the effect of the use upon the potential market for or value of the copyrighted work.
AI models are having a massive effect on the market for and value of the works they train on, as is being widely discussed in multiple industries. Art, writing, voice acting, code; in any market AI can generate content for, the value of such content goes down. (This argument does not require AI to be as good as humans. Even if the high end of every market produces work substantially better than AI, flooding a market with unlimited amounts of cheap/free low-end content still has a massive effect.)
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#138Earlier quoted context omitted.
How is it different from asking to me to summarize anything? I could have bought the book, or read the Wikipedia page, or listened people talking about it, or downloaded the torrent. In all those cases my summary could be right or could be wrong. If the rights holders know that I dowloaded the torrent they could sue me. In the other cases they can't. What if it turns out that OpenAI bought a copy of every book ingest…
I feel like thats one of the many questions regulators and law makers are going to be asked long term. I'm sure buying the book for "commercial purposes" like that would't be appropriate, but then again, does that mean if I read it and then summarize it in my work, or regurgitate its info as part of my job...I'm violating a license? A world where humans have special permissions but LLMs don't seems pretty interesting…
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#139Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#140> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…
I for one am quite happy that AI folks are basically treating copyright as not existing. I strongly hope that the courts find that LLM weights, and the datasets are "fair-use" or whatever other silly legal justification. Aaron Swartz was a saint.
Five years from now, will OpenAI actually be open, or will it be a rent seeking org chasing the next quarterly gains? I expect the latter.