Live data from Hacker News

A federal judge sides with Anthropic in lawsuit over training AI on books

techcrunch.com

181–190 of 222 posts

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#181
post #179
post #64

Earlier quoted context omitted.

> Worse, they’re using it for massive commercial gain, without paying a dime upstream to the supply chain that made it possible. If there is any purpose of copyright at all, it’s to prevent making money from someone’s else’s intellectual work. This makes no sense. If I buy and read a book on software engineering, and then use that knowledge to start a career, do I owe the author a percentage of my lifetime earnings?…

> If I buy and read a book on software engineering You're comparing that you as an individual purchase one copy of a book to a multi-billion dollar company systematically ingesting them for profit without any compensation, let alone proportional? > do I owe the author a percentage of my lifetime earnings? No, but you are a human being. You have a completely different set of rights from a corporation, or a machine. Fo…

Does copyright law apply differently to humans Vs organisations?

> without any compensation,

Didn't Anthropic buy the books?

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#182

Humans read books. AI/LLMs do not read. I think there's an inherent difference here. If the LLM is making a copy of the entire book in it's memory, is that copyright infringement? I don't know the answer to that, but it feels like Alsup is considering this fair use argument in the context of a human, but it's nothing like a human and needs to be treated differently.

LLMs do not "make a copy of the entire book in its memory" so that specific question is kind of moot.

Its already established it can recite whole Hairy Potter and Carmacks Fast Inverse word for word. Just because it uses fancy compression doesnt mean its not a copy.

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#183

The US legal systel is bending over backwards to help AI development. The arguments border on nonsense.

Can you offer some examples from this ruling? It seems pretty reasonable on a first read.

Judge decided having an output filter on your AI makes it ok for it to contain full copy of copyrighted work.

Its like saying it should be legal for me to have this Judges nudes obtained 100% illegally as long as I pixelate all the naughty bits.

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#184
post #174

Earlier quoted context omitted.

> Says who? Artists. https://en.wikipedia.org/wiki/SAG-AFTRA > How on earth are those things mutually exclusive? Put those on a spectrum and rethink what I said. > completely irrelevant to whether or not it is copyright infringement _Again_, leave aside law minutiae and hypotheticals.

> > Says who? > Artists. > https://en.wikipedia.org/wiki/SAG-AFTRA Do you have a link that has their stance on how AI is harming culture? The best I could find is https://www.sagaftra.org/contracts-industry-resources/member... I can't find anything in there or its linked articles about culture. I do find quite a bit about synthetic performers and digital replicas and making sure that people who do voice acting don't…

> they want to make sure that people are in control of it and people are compensated for the works that are created

Nice! Now you just need to connect the dots from your own conclusion to my initial statement.

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#185

One aspect of this ruling [1] that I find concerning: on pages 7 and 11-12, it concedes that the LLM does substantially "memorize" copyrighted works, but rules that this doesn't violate the author's copyright because Anthropic has server-side filtering to avoid reproducing memorized text. (Alsup compares this to Google Books, which has server-side searchable full-text copies of copyrighted books, but only allows user…

No. You are free to memorize any copyrighted work. You are just not free to distribute it.

The model itself does not constitute a copy. Its intention is clearly not to reproduce verbatim texts. There would be far cheaper and infinitly more accurate ways to do that if that was the goal.

Appart from the legalities, it would be horrifying if copyright reached into the AI realm to completely styfle progress for, lets be honest, mainly the profits of a few major IP corporations.

I do however understand some creatives are worried about revenue, just like the rest of us. But just like the rest of us, they to live in a world that can only exist because 99.99% of what it took to build that world was automated or tool enhanced, impacting someone's previous employment or business.

We are in a world of unprecedented change, only to be immediatly supassed by the next day's rate of change. This both scares and fascinates me.

But that change and its benefits being held only in the bowels of corporate/government symbiotic entities would scare me a hell of a lott more. Open Source/weights is the only way to have a small chance to keep this at bay.

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#186

Earlier quoted context omitted.

I mean the human brain can memorize things as well and it’s not illegal. It’s only illegal if said memorized thing is distributed.

Because humans have rights AI models do not.

They use to say the same thing about black people.

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#187

Earlier quoted context omitted.

This is nonsense, in my opinion. You aren't "hearing" anything. You are literally creating a work, in this case, the model, derived from another work. People need to stop anthropomorphizing neural networks. It's a software and a software is a tool and a tool is used by a human.

It is easy to dismiss, but the burden of proof would be on the plaintiff to prove that training a model is substantially different than the human mind. Good luck with that.

That makes no sense as a default assumption. It's like saying FSD is like a human driver. If it's a person, why doesn't it represent itself in court? What wages is it being paid? What are the labor rights of AI? How is it that the AI is only human-like when it's legally convenient?

What makes far more sense is saying that someone, a human being, took copyrighted data and fed it into a program that produces variations of the data it was fed. This is no different from a photoshop filter, and nobody would ever need to argue in court that a photoshop filter is not a human being.

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#188
post #168

Earlier quoted context omitted.

The "fairness" argument is weaker than the "sustainable creation" one. If LLMs could create quality literature, or social media create in-depth reporting, then I'd have no problem with the tide of technological progress flowing. Unfortunately, recent history has shown that it's trivial for the market to cannibalize the financial model of creators without replacing it . And as a result, society gets {no more that thin…

I guess the optimistic take would be that we will get novel, insightful synthesis of disparate fields of knowledge that no human so far was ever able to hold in their mind to contemplate their interrelations. And this will elevate the human spirit etc. The equivalent to the take that the Internet will bring peoples together and foster better understanding and love between people who so far were not in dialogue and th…

I suppose another strongman would be that LLMs substantially decrease the cost of human creation (i.e. the HITL assistant use case) while producing an output of equivalent quality.

As a result of this, everything gets cheaper and more plentiful.

The counterargument I'd make to that would be the requirement that the human have creative skills, which might atrophy in the absence of business models supporting a career creating.

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#189
post #180
post #166

Earlier quoted context omitted.

Wouldn't a model that can recite training data verbatim be larger than necessary? Exact text isn't coming from nowhere, no matter how efficiently the bits are encoded, and the same effectiveness should be achievable by compressing those portions of the model.

Maybe we are all just LLMs. If the books were written by a language producing algorithm in a human mind, maybe there’s not as much raw data there as it seems, and the total information can in fact be stored in a surprisingly small set of weights.

I imagine it's not inconceivable that at very high dimensions and with the right architectures stochastic compression can be unexpectedly efficient. It would be strange if the end result of AI research is realizing we're solving a compression problem (and that our brains do too).

Re: A federal judge sides with Anthropic in lawsuit over training AI on books

#190
The reason I made books3 was to help force a decision on this issue. I’m happy to see that it’s settled, and that it’s legal for robots to read books.

It’s also proof that an individual scientist can still change the world, in some small way. Believe in yourself and just focus on your work, even if the work is controversial.

(I’m late to the thread, so ~nobody will see this. But it’s the culmination of about five years of work for me, so I wanted to post a small celebratory comment anyway. Thank you to everyone who was supportive, and who kept an open mind. Lots of people chose to throw verbal harassment my way, even offline, but the HN community has always been nice.)

Post reply on HN