Live data from Hacker News

A federal judge sides with Anthropic in lawsuit over training AI on books

techcrunch.com

201–210 of 222 posts

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#201
post #57

Earlier quoted context omitted.

While no one wants anyone to steal a car, almost no one would mind freely cloning a car. The trouble truly is that 3d-printing hasn't gotten that good yet.

The car would be unlikely to exist if its maker had to expect free clones without compensation. So yes, people would mind.

Completely untrue. If some clever engineer or consortium of engineers designed a 3D-printable car for 3D printing-and-manufacturing companies to make then it surely would exist. If you buy one from a Ford dealership you'd be getting the Ford-branded version which may have their own tweaks to the design.

It makes perfect sense to me that the big carmakers could get together some day and develop a handful of car platforms that all their cars will be built upon. That way they can buy the parts from any number of manufacturers (on-demand!) and save themselves a ton of money.

They kind of already do that, actually =)

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#202
post #6

Earlier quoted context omitted.

Anthropic won't submit a spreadsheet of all the books and whether they were purchases or not. So trivially, not every book stolen is shown to be later purchased. As just a matter of society, I don't think you want people say stealing a car and then coming back a month later with the money.

Stealing a car deprives the previous owner of the car of possession and use. It is a criminal charge and you will be punished for it regardless of the monetary value of the car. The owner of the car could also sue the thief for financial damages caused by not having the car for a month, which won't be more than the cost of an equivalent rental for a month, so it's not even worth bothering. Copyright infringement does…

Your take on how copyright infringement works only counts for unregistered copyrights. If the copyrighted works are registered with the copyright office statutory damages apply:

https://www.law.cornell.edu/uscode/text/17/504

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#203
post #3

The HN crowd dislikes brick-and-mortar landlords but often sides with charging rent for certain bits. Which side will prevail? Interesting excerpt: > “We will have a trial on the pirated copies used to create Anthropic’s central library and the resulting damages,” Judge Alsup wrote in the decision. “That Anthropic later bought a copy of a book it earlier stole off the internet will not absolve it of liability for the…

>If they did realize a mistake and purchased copies after the fact, why should that be insufficient? 1. You're assuming this was some good faith "they didn't know they were stealing" factor. They use someone else's product's for commercial use. I'm not so charitable in my interpretation. 2. I'm not absolved of theft just because I go back and put money on the register. I still sttole, intentionally or not

Google trained their AI on stuff they scraped without knowing whether it was pirated content. Why should it be different for Anthropic?

Google literally scrapes pirated content all day every day. When they do that they have no idea if the content was legally placed on that website. Yet, they scan and index it anyway because there's actually no way to know (at all!). There's no great big database of all copyrighted works they can reference.

I'm not saying Meta and Anthropic didn't know they were pirating content. I'm saying that it should be moot since they never distributed it. You can't claim a violation of copyright for content that was never actually "copied" (aka distributed). The site/seeders that uploaded the content to Meta/Anthropic are the violators since copyright is all about distribution rights.

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#204
post #182

Earlier quoted context omitted.

LLMs do not "make a copy of the entire book in its memory" so that specific question is kind of moot.

Its already established it can recite whole Hairy Potter and Carmacks Fast Inverse word for word. Just because it uses fancy compression doesnt mean its not a copy.

It can recite something like 80% of Harry Potter with carefully crafted prompts. If you take half a sentence from Harry Potter then tell the LLM to predict the rest it will complete it. That's what they did in that study you're referring to.

It's not even remotely the same thing as "can recite whole Harry Potter." If you ask an LLM to regurgitate Harry Potter it won't be able to do so because that's not how they work. They're prediction engines and it just so happens that Harry Potter quotes/excerpts are so pervasive on the Internet that the LLMs ingress ranks that style of wording higher than other styles.

Ask it to regurgitate some other, less-popular work. Do it for hundreds or thousands of them. You'll quickly find that those two examples you gave are the exceptions and that LLMs can't pull it off. They won't even get close.

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#205

Earlier quoted context omitted.

As you point out, people make rules ("laws") which benefit them . I care about fairness and justice though, even if I am a minority. Fundamentally, fair compensation is based on the amount of work put in (obviously taking skill/competence into account but the differences between people in most disciplines probably don't span a single order of magnitude, let alone several). The ultimate goal should be to prevent peopl…

>Fundamentally, fair compensation is based on the amount of work put in. I think there is a problem with your initial position. Nobody is entitled to compensation for simply working on something. You have to work on things that people need or want. There is no such thing "fair compensation". It is "unfair" to take the work of somebody else and sell it as your own. (I don't think the LLMs are doing this.)

Yes, I meant when working on the same thing (which has a specific value as a whole).

If the LLM and its output are based on 10^12 hours of work, out of which 10^6 is working on the code of the LLM itself and 10^12-10^6 (so roughly still 10^12) is working on the training data, does it make sense for only those working on the 10^6 to be compensated for the work?

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#206
post #183

Earlier quoted context omitted.

Judge decided having an output filter on your AI makes it ok for it to contain full copy of copyrighted work. Its like saying it should be legal for me to have this Judges nudes obtained 100% illegally as long as I pixelate all the naughty bits.

Full ruling is here ( https://storage.courtlistener.com/recap/gov.uscourts.cand.43... ) The analogy the judge gives is to how Google Books walked the tightrope on copyright: they maintain an archive of all the books for indexing and search purposes, and can display excerpts to help you confirm that's what you're looking for. The excerpts are constrained so you can't read the whole book by scanning the excerpts. If po…

> Anthropic is liable for that copying

That's yet to be determined. The judge ruled that an entirely separate trial will be necessary to determine if Anthropic violated specific copyrights when they downloaded books from pirate websites and what the damages would be if they did so.

So far no court case has ruled downloading to be a violation of copyright. In Sony BMG Music Entertainment v. Tenenbaum and Capitol Records, Inc. v. Thomas-Rasset the courts ruled that downloading and then sharing the content constituted a violation of copyright law. Those are the only two cases I'm aware of where a ruling was made (relevant to this).

The courts need to be very careful with any such ruling because search engines download pirated content all day every day. If the mere act of downloading it violated copyright law then that will break the Internet (as we know it).

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#207

> “We will have a trial on the pirated copies used to create Anthropic’s central library and the resulting damages,” Judge Alsup wrote in the decision. “That Anthropic later bought a copy of a book it earlier stole off the internet will not absolve it of liability for theft but it may affect the extent of statutory damages.” I'm not sure why this alone is considered a separate issue from training the AI with books. B…

> Keep in mind, you can legally engineer EULAs in such a way that merely purchasing the work surrenders all of your fair use rights.

That has yet to be determined in a court of law. Just like: You can write a contract to kill but that won't make it legal.

The Supreme Court ruled that Fair use is an essential component that makes copyright law compatible with the First Amendment. I highly suspect that if if ever comes up in the SCOTUS they will rule that only signed contracts can override Fair Use. Meaning: Clickwrap agreements or broad contracts required by ebook publishers (e.g. when you use their apps) don't count.

Also, if you violate a contract by posting an excerpt of an ebook you purchased online would require the publisher to sue you in court (or at least force arbitration) over that contract violation. They could not use tools like the DMCA in such instance to enforce a takedown request.

There's no, "Hey! They're violating our contract, I swear!" takedown feature in contract law like there is with copyright law (the DMCA).

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#208

One aspect of this ruling [1] that I find concerning: on pages 7 and 11-12, it concedes that the LLM does substantially "memorize" copyrighted works, but rules that this doesn't violate the author's copyright because Anthropic has server-side filtering to avoid reproducing memorized text. (Alsup compares this to Google Books, which has server-side searchable full-text copies of copyrighted books, but only allows user…

I am yet to have anyone explain to my why LLM memorisation is worse than Google images or a similar service caching thumbnails for faster image searches. Or caching blurbs of news stories for faster reproduction at search time.

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#209
post #10
post #2

Broadly summarizing. This is OK and fair use: Training LLMs on copyrighted work, since it's transformative. This is not OK and not fair use: pirating data, or creating a big repository of pirated data that isn't necessarily for AI training. Overall seems like a pretty reasonable ruling?

But those training the LLMs are still using the works, and not just to discuss them, which I think is the point of fair use doctrine. I guess I fail to see how it's any different from me using it in some other way? If I wanted to write a play very loosely inspired by Blood Meridian, it might be transformative, but that doesn't justify me pirating the book. I tend to think copyright should be extremely limited compare…

If I buy a book, and use it to prop up the table on which I build a door, I dont owe the author any additional money over what I paid for it.

If I buy a book, and as long as the product the book teaches me to build isnt a competing book, the original author should have no avenue for complaint.

People are really getting hung up on the computer reading the data and computing other data with it. It shouldnt even need to get to fair use. Its so obviously none of the authors business well before fair use.

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#210
post #2

Broadly summarizing. This is OK and fair use: Training LLMs on copyrighted work, since it's transformative. This is not OK and not fair use: pirating data, or creating a big repository of pirated data that isn't necessarily for AI training. Overall seems like a pretty reasonable ruling?

If a publisher adds a "no AI training" clause to their contracts, does this ruling render it invalid?

Fair Use and similar protections are there to protect the end user from predatory IP holders.

First, I dont think publishers of physical books in the US get the right to establish a contract. The book can be resold for instance and that right cannot be diminished. But secondly adding more cruft to the distribution of something that the end user has a right to transform, isn't going to diminish that right.

Post reply on HN