Live data from Hacker News

Sarah Silverman is suing OpenAI and Meta for copyright infringement

theverge.com

261–270 of 599 posts

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#261

Earlier quoted context omitted.

It's been quite a few years, but AFAIK he didn't hack JSTOR. He downloaded papers en mass using a guest MIT account that had legal access to JSTOR. Maybe he violated their terms, but that is not illegal. He did illegally trespass an unlocked MIT switch closet to do this. They blocked several IPs but his script would rotate to continue. The downloading was over a week or two, enough for security to set up a camera in…

> AFAIK he didn't hack JSTOR. He downloaded papers en mass using a guest MIT account that had legal access to JSTOR. Potato, potahto. Or, like kids these days say it, "corporate wants you to find differences between these two pictures...". Fact is, from the POV of the legal system, "using a guest account that had legal access to" a system, but to which (the account) you didn't have legal access, would typically be se…

> Fact is, from the POV of the legal system, "using a guest account that had legal access to" a system, but to which (the account) you didn't have legal access

I don't think that is accurate.

He had legal access to the account. The account had legal access to the service.

The argument was that downloading articles en masse was an _abuse_ of the service, which was a violation of the Terms of Service and therefore a CRIMINAL ACT.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#262

> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…

Let's take a second to remember that copyright is the reason ~every child doesn't have access to ~every book ever written.

While it might be too disruptive to eliminate copyright overnight, we should remember that our world will be much better and improve much faster to the extent we can reduce copyright's impact.

And we should cheer it on when it happens. A majority of the world's population in 2023 has a smartphone. Imagine a world where a majority of the world had access to every book ever digitized, and could raise their children on these books!

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#263
On one hand: should I pay copyright owners for the privilege of sharing what I learned, regardless of how I obtained it?

On the other: I am not supposed to share sensitive data, confidential trade secrets, or privileged national security intelligence that I learned without the necessary authorization.

Applied to an LLM, how would that work?

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#264

Earlier quoted context omitted.

In general copyright (in the us) doesn't cover transformational usage. If you can argue that the nature of your use is transformative you might be good.

"The transformative use concept arose from a 1994 decision by the U.S. Supreme Court. In Campbell v. Acuff-Rose Music, the Court focused not only on the small quantity taken from the copyrighted work but also on the transformative nature of the defendant’s use. The case concerned a song by the group 2 Live Crew entitled "Pretty Woman," which, according to an affidavit, was meant to "through comical lyrics, satirize t…

[deleted]

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#265

Earlier quoted context omitted.

That's not was he's saying at all. He's saying you can train an AI on copyrighted material just like people can learn from copyrighted material. If you acquire the material illegally that a separate issue that training AI doesn't give you any protection against.

>you can train an AI on copyrighted material just like people can learn from copyrighted material. No amount of whining and hand wringing from engineers will ever make this true. This is for the courts to decide. A reasonable interpretation, in my eyes, is that the training process is a black box which takes in copyrighted works and produces a training model. The training model is a derivative work of the inputs. It…

[flagged]

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#266
post #23

Earlier quoted context omitted.

... on a book that they’d illegally acquired then read.

But don't you see what a strange argument that is? It doesn't matter, the student did nothing wrong, I don't want to live on a planet where we put DRM into peoples brains (or AI for that matter) to enforce this absurd and overreaching idea of intellectual property. And besides the publishers extorting thousands from young students forced to buy their overpriced mediocre textbooks, warrants any copyright infringement…

If the book was acquired illegally, the entity that suffered the loss may have a claim for the illegal acquisition. Meta and OpenAI have the money to buy a copy of every book under copyright that they have their AI read for training. I have more sympathy for losses suffered by a living person that produced creative works than I do for textbook mills. I also have sympathy for open source software authors that applied their creativity to create source code that is spit out verbatim by Copilot without adhering to license terms.

I see that Thomas’ Calculus is up to the 14th edition, priced at about 20x the hourly wage that a college student will earn. Thomas and Finney Calculus editions 6 - 8 are on shelves or in boxes somewhere in my house. Each of those cost me or my wife about 20x the student hourly wage back in the day. I bet calculus hasn’t changed a lot in the past 30 years to justify all of these editions. I blame universities for allowing this industry to thrive.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#267

Earlier quoted context omitted.

Possession of copyrighted material without permission of the creator is illegal, but yes -- rights holders don't really go after infringers except maybe via ISP three strikes crap. They're very much incentivized to change their behavior for AI scraping, though.

I am not a lawyer, but my understanding is that copyright law typically regulates the unauthorized reproduction, distribution, public display, or creation of derivative works of copyrighted materials. Possession of copyrighted material in itself is not illegal. It's how you use that material that could potentially violate copyright laws.

If we are speaking of US criminal law, then what you say is true. It doesn't mean that is true in other regions, nor does it mean you wouldn't be held liable in a civil court.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#268

Earlier quoted context omitted.

Isn’t the burden of proof on the other side?

Not when OpenAI publicly declared they trained on pirated works. I can’t imagine “we can’t tell if this is the result of the illegal thing we did or not” is going to stand up very well, nor does it bode well for any refutation of the plaintiff’s depiction of their intent. Part of fair use consideration is commercial impact and when you steal a bunch of books to train your AI model, it’s hard to refute that the impact…

… it’s hard to refute that the impact is not negative or that you didn’t intend commercial harm.

So this is another thing I don’t understand. Is the claim that fewer people will buy Silverman’s book because ChatGPT is able to provide a summary? If so, call me skeptical.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#269

Earlier quoted context omitted.

You can currently download any copyright work from any pirate site and be free of all consequences, and this has always been so. You just can't upload, since that counts as distribution, triggering civil and criminal penalties written in an age before the Internet when only shady commercial operators would distribute unlicensed copyrighted works.

> You can currently download any copyright work from any pirate site and be free of all consequences, and this has always been so. No. By virtue of "download" of a file, you are making a copy of it which is in violation of US copyright (and lots of countries. You're unlikely to be sued or prosecuted for it, but that doesn't make it legal.

This is just completely untrue.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#270

Earlier quoted context omitted.

Possession of copyrighted material without permission of the creator is illegal, but yes -- rights holders don't really go after infringers except maybe via ISP three strikes crap. They're very much incentivized to change their behavior for AI scraping, though.

I am not a lawyer, but my understanding is that copyright law typically regulates the unauthorized reproduction, distribution, public display, or creation of derivative works of copyrighted materials. Possession of copyrighted material in itself is not illegal. It's how you use that material that could potentially violate copyright laws.

Yet.. If you configure your torrent client to never seed back (i.e. never upload or transmit content to others) and download from a public tracker, when the copyright infringement notice comes from your ISP, this explanation will change nothing.

Notably, I think this is wrong - as per the legal definition, publishers, ISPs, and courts should only hold you accountable if you helped distribute via uploading.

Post reply on HN