Live data from Hacker News

Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

apnews.com

511–520 of 654 posts

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#511
post #293

Earlier quoted context omitted.

This settlement has basically nothing to do with LLMs. At least not as far as the courts are concerned. Alsup ruled [0] that feeding a book into an LLM is transformative and counts as fair use. Especially when they purchased a physical copy of the book, scanned it, and destroyed the original. But if I'm reading the ruling correctly, Anthropic might have been fine even with feeding pirated books into their LLM (as lon…

As far as I'm concerned, the courts are wrong, and training on ill gotten copyrighted material is not fair use. Given the clear value of highly trained LLMs, the investment they have taken on, and the amount of disruption to the existing economy they stand to make, in a just world, the people who created the training data deserve some level of compensation. I think, in the US, they are very afraid of falling behind C…

Adobe’s ereaders had a disclaimer that their books cannot be read aloud. There’s clearly precedent that this sort of transformation was disallowed by publishers at the time. Interestingly, at least the audiobook of the latest dungeon crawler Carl has a disclaimer that it can’t be used to train AI

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#513
post #404

Earlier quoted context omitted.

I'll link to a previous comment of mine: https://news.ycombinator.com/item?id=48968156 > You need to have a very specific and 'creative' / 'substantial' expression of an idea for copyright to apply. The output of an LLM can be easily be such, but usually not. This is incomplete with current US law. You need the above (the typical copyright qualifiers) AND evidence of substantial human involvement in the creation. Min…

Just to be clear, what you're referring to is the current US standard for whether a work is copywritable, not whether training on data and "regurgitating existing ideas" is fair-use. The latter is what the GP comment was about: > There needs to be a royalty payment based on if the AI regurgitates existing ideas. That is probably the correct way to legislate this. If anything a human does can instantly be copied by an…

Correct. I just wanted to clarify the statement that parent made, as it seems like lots of people have a misassumption about the copyrightability of autonomous in the United States.

Expect it will be clarified and/or changed by law given how much money is at stake, but the current state is what the current state is.

If I were developing key IP with agents, I'd be very careful to document my human contribution.

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#515
post #420

Earlier quoted context omitted.

How many times were hugely popular books rejected before a publisher decided they were worthy? Copyright far more protects the wealthy than the good. They don't need to sell your book, they just need to own the book that people are buying right now. Giving your book a chance to sell would dectract from those sales. If there were no copyright anyone trying to sell the book $1 cheaper would be undercut by someone selli…

Well, there is nothing to distribute if the author is not incentivized to write... which you seemed to skip past.

If money is the only incentive, then it's not a product of artistic work.

Also current copyright laws only exists to fulfill the constitutional mandate to promote the progress of science and useful arts. There are a lot of alternative ways to fulfill that mandate that don't include a lot of the baggage we have presently in copyright law which is now slowing down progress.

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#516
post #461
post #248

Earlier quoted context omitted.

You assume your premise. But plenty of "intellectual work" is already done without legal cover. It just typically attracts normal profits, rather than super-normal rent-seeking ones.

I, too, hate rentseeking. Owning one's intellectual output, however, is not in and of itself rentseeking.

Rents are just amounts beyond what's needed to cause the thing to exist. At the point of copying something, it already exists.

So payments for the right to do so aren't payments required to bring anything new into existence at that point, save for the legal fiction.

Now you might argue that the future copy-licencing rents are necessary to bring the _original_ creation into being. But that doesn't make them _not rents_.

But I would say that's the second assumption you're baking in here.

As in, we live in a world where e.g. the movie Toy Story exists. Now, certainly Toy Story does provide some good or value to the world. But I don't think you can assume such things provide more value than e.g. open science, free transformation of works, etc.

I get that people enjoy our current IP culture but saying certain things wouldn't exist in an IP-free world is just an argument from consequences that doesn't even really compare consequences between the two.

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#517
Now view this in contrast with what happened to Aaron Swartz

> According to state and federal authorities, Swartz used JSTOR, a digital repository,[79] to download a large number[note 2] of academic journal articles through MIT's computer network over the course of a few weeks

> ...federal prosecutors filed a superseding indictment adding nine more felony counts, increasing Swartz's maximum criminal exposure to 50 years of imprisonment

> ...On the evening of January 11, 2013, Swartz's girlfriend, Stinebrickner-Kauffman, found him dead in his Brooklyn apartment.[80][116][117] A spokesperson for New York's Medical Examiner reported that he had hanged himself

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#518
post #493
post #453

Earlier quoted context omitted.

Whatever "training" is, if you can't persuade the machine to spit substantially the same text back out verbatim, it's clearly not something that falls under copy right law either, because there's no copy. Yes, for some texts that's possible. But for the vast majority, it is not.

> spitting out verbatim text The New York Times lawsuit is resting on the point that large chunks of undigested articles can be vomited out. OpenAI tried to have the lawsuit thrown out but the courts permitted it to continue. The Times... alleged that OpenAI's ChatGPT and Microsoft's Copilot had produced near-verbatim replicas of copyrighted articles, that the chatbots generated hallucinated content falsely attribute…

> spitting out verbatim text

> had produced near-verbatim replicas

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#519
post #452

Earlier quoted context omitted.

Yes, ultimately the problem is that the law is vague or inadequate. The courts have their definitions of fair use, which are their best efforts at interpreting the law, and I have mine, which is different.

William Roper: "So, now you give the Devil the benefit of law!" Sir Thomas More: "Yes! What would you do? Cut a great road through the law to get after the Devil?" William Roper: "Yes, I’d cut down every law in England to do that!" Sir Thomas More: "Oh? And when the last law was down, and the Devil turned ’round on you, where would you hide, Roper, the laws all being flat? This country is planted thick with laws, fro…

This is why the idea of being "Vogelfrei" or "lawless" was honestly a terrifying concept in the middle ages. They are neither bound by law, nor protected by law.

A lawless man can be struck down with force without persecution by law, because they are lawless.

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#520
post #97

Judge Alsup issued the original order that determined they were liable for piracy but that training LLMs on books was fair use. It's worth reading if you're interested in the topic. https://www.courtlistener.com/docket/69058235/231/bartz-v-an...

Alsup is an interesting judge. He has handled several important tech cases, such as Oracle v Google, and Waymo v Uber. He's also a longtime hobbyist programmer working in BASIC, much of it in support of his ham radio hobby. Screenshots of his shortwave propagation prediction program here [1]. [1] https://www.theverge.com/2017/10/19/16503076/oracle-vs-googl...

He was one of the few judges that understood tech. Unfortunately he retired last year.
Post reply on HN