Live data from Hacker News

Anthropic agrees to pay $1.5B to settle lawsuit with book authors

nytimes.com

181–190 of 761 posts

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#181
post #153
post #139

Earlier quoted context omitted.

There’s alternatives to wiping out the company that could be fair. For example, a judgement resulting in a shares of the company or revenue shares in the future rather than a one time pay off. Writers were the true “foundational” piece of LLMs, anyway.

If this is an economist idea of fair, where is the market? If someone breaks into my house and steals my valuables, without my consent, then giving me stock in their burglary business isn't much of a deterrent to them and other burglars. Deterrence/prevention is my real goal, not the possibly of a token settlement from whatever bastard rips me off. We need the analogue of laws and police, or the analogue of homeowner…

I don't much like the idea of settling in stock, but I also think you're looking for criminal law here. Civil law, and this is a civil suit, is far more concerned with making damaged parties whole than acting as a deterrent.

I understand that intentional copyright infringement is a crime in the US, you just need to convince the DOJ to prosecute Anthropic for it...

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#182

Earlier quoted context omitted.

> Buying used copies of books, scanning them, and training on it is fine. Buying used copies of books, scanning them, and printing them and selling them: not fair use Buying used copies of books, scanning them, and making merchandise and selling it: not fair use The idea that training models is considered fair use just because you bought the work is naive. Fair use is not a law to leave open usage as long as it doesn…

Buying used copies of books, scanning them, training an employee with the scans: fair use. Unless legislation changes, model training is pretty much analogous to that. Now of course if the employee in question - or the LLM - regurgitates a copyrighted piece verbatim, that is a violation and would be treated accordingly in either case.

> Buying used copies of books, scanning them, training an employee with the scans: fair use.

Does this still hold true if multiple employees are "trained" from scanned copies at the same time?

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#183
post #67

One thing that comes to mind is... Is there a way to make your content on the web "licensed" in a way where it is only free for human consumption? I.e. effectively making the use of AI crawlers pirating, thus subject to the same kind of penalties here?

I'd argue you don't actually want this! You're suggesting companies should be able to make web scraping illegal. That curl script you use to automate some task could become infringing.

>I'd argue you don't actually want this! You're suggesting companies should be able to make web scraping illegal.

At this point, we do need some laws regulating excessive scraping. We can't have the ineternet grind to a halt over everyone trying to drain it of information.

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#184

Earlier quoted context omitted.

I think he implies that because one can borrow hypothetically any book for free from a library, one could use them for legal training purposes, so the requirement of having your own copy should be moot

Libraries aren’t just anarchist free for alls they are operating under licensing terms. Google had a big squabble with the university of Illinois Urbana Champaign research library before finally getting permission to scan the books there. Guess what, Google has the full text but books.google.com only shows previews, why is an exercise to the reader literally

Libraries are neither anarchist free for alls nor are they operating under licensing terms with regards to physical books.

They're merely doing what anyone is allowed to with the books that they own, loaning them out, because copyright law doesn't prohibit that, so no license is needed.

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#185

Earlier quoted context omitted.

Google scanned many books quite a while ago, probably way more than LibGen. Are they good to use them for training?

They litigated this a while ago and my understanding was that they were able to claim fair use, but I'm no expert. What I'm wondering is if they, or others, have trained models on pirated content that has flowed through their networks?

Books.Google.Com was deemed fair use because it only shows previews, not full downloads. Internet Archive is still under litigation iirc besides having owned a physical copy of every book they ever scanned (and keeping a copy in their warehouses) they let people read the whole thing.

I’m surprised Google hasn’t hit its competitors harder with the fact that they actually got permission to scan books from its partner libraries and Facebook and OpenAI just torrented books2/books3, but I guess they have aligned incentive to benefit from a legal framework that doesn’t look to closely at how you went about collecting source material

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#186
post #137

Earlier quoted context omitted.

I feel like proportionality is related also to the scale. If a student pirates a textbook, I’d agree that 100x is excessive, but this is a corporation handsomely profiting off of mass piracy. It’s crazy to imagine, but there was surely a document or slack message thread discussing where to get thousands of books, and they just decided to pirate them and that was OK. This was entirely a decision based on ease or cost,…

> handsomely profiting Well actively generating revenue at least. Profits are still hard to come by.

Operating profits certainly but if you include investments the big players are raking it in aren't they?

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#187
post #63

Anyone have a link to the class action? I published a book and would love to know if I'm in the class.

Docket: https://www.courtlistener.com/docket/69058235/bartz-v-anthro...

Proposed settlement: https://storage.courtlistener.com/recap/gov.uscourts.cand.43...

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#188
post #8

To be very clear on this point - this is not related to model training. It’s important in the fair use assessment to understand that the training itself is fair use, but the pirating of the books is the issue at hand here, and is what Anthropic “whoopsied” into in acquiring the training data. Buying used copies of books, scanning them, and training on it is fine. Rainbows End was prescient in many ways.

I wonder what Aaron Swartz would think if he lived to see the era of libgen.

Didn't he get in trouble for contributing to sci-hub before he died?

Re: Anthropic agrees to pay $1.5B to settle lawsuit with book authors

#190
post #174

Earlier quoted context omitted.

Why would they earn more from models reading their works than I would pay to read it?

Because ones doing the training are profiting from it. Ai is not a human with limited time. And it is also owned by a company not a legal person. I might find argument of comparing it to human when it is fully legal person and cutting power to it or deleting is treated as murder. Before that it is just bullshit. And fundamentally reason for copy right to exist is to support creators and to promote them to create more…

If I buy a book, learn something, and then profit from it, should I also be paying more than the original price to read the book?

> Ai is not a human with limited time

AI is also bound by time, physics, and limited capacity. It does certain things better or faster than us, it fails miserably at certain things we don't even think about being complex (like opening a door)

> And it is also owned by a company not a legal person.

For the purpose of legalities, companies and persons are relatively equivalent, regardless of the merits, it is how it is

> In world where massively funded companies can freely exploit their work and even in many case fully substitute that principle is failed.

They paid for the books after getting caught, the other companies are paying for the copyrighted training materials

Post reply on HN