Live data from Hacker News

Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

apnews.com

371–380 of 654 posts

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#371
post #340

Earlier quoted context omitted.

So now that we have a magical paraphrasing machine, we can just run any copyrighted work through it to remove the copyright? Cool, I get a GPL version of Microsoft Office.

That’s what a brain is

Maybe but your brain is not running 24/7 capable of outputting thousand if not millions of tokens per hour, all while having ingested nearly the entire internet.

If yours do that, maybe we can redefine what copyrighting and patenting means for humans

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#372
post #299

Earlier quoted context omitted.

What you are saying leads to "pulling the ladder behind you" effect on creativity. It's impossible to protect more than substantial similarity and still allow creativity to exist.

If a human makes something there should be broad protection for creativity, if a LLM generates something there should be extremely limited protection for creativity. You should not be able to mass generate images in a particular artists style and claim it as fair use, even if a human making the same images would have protection.

But the human made the LLM. An LLM is categorically “I built a thing that built a thing” and if the output of that category has no protections then all automation and ‘machine at the final step’ is in trouble.

What about aleatory music (music left at least partially to chance)? Or Autechre - they have whole albums and live performances built on automation software. They built the logic and added randomization, necessarily removing themselves from the final output.

Is spin art not copyrightable? If I build a simple machine that spins paper, then do no more than drop paint on it, the result is not mine to copyright? I didn’t choose the output, I merely built the machine and the rest was created by pure chance. “But you chose the paint” - and if I didn’t? What if my art uses AI to perform sentiment analysis on the top news articles of the day and it drops colors matching the emotional tone of the news onto the spin art machine. I have no control over it and the output is machine generated, but is the result not just the final step of an entire process I created? Was the result of the creative idea not part of the creativity itself?

If I build an automated laboratory to test every combination of a problem space, is a resulting success not patentable? What if the problem is too large to permute, so I added a random selection process to it? I’m not even controlling what’s being tested, but if it finds success is that not my contribution? The machines did the work, the selection was random, there was no human in the loop; what then?

The internals of an LLM may be mysterious to some, but I assure you it’s just fixed automation with a random number generator sometimes tacked onto it, but randomization is optional too.

I built that LLM. I decided what text to input for training, I curated the information, I wrote the algorithm, I decided the layers and hyper-parameters, I decided the RLHF pairs to train, then I put a few drops of paint from my bottle of language into the automated machine. I decided and built every single step of the system, but that output is not part of my process? If I pipe the LLM text output to a paint dispenser hovering over paper, set to squeeze out drops based on syllables, would you protect my artwork then?

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#373
post #271

A one time payment like 1.5B doesn’t do anything. There needs to be a royalty payment based on if the AI regurgitates existing ideas. That is probably the correct way to legislate this. If anything a human does can instantly be copied by an LLM, and then sent to all its subscribers, things need to change

Ideas are not protected by copyright, nor are facts. You need to have a very specific and 'creative' / 'substantial' expression of an idea for copyright to apply. The output of an LLM can be easily be such, but usually not.

> … very specific …

That phrase is doing a lot of work. In the US, any writing is automatically protected by copyright. (This comment, for example.) Whether the author can claim infringement is a can of worms: legal costs, fair use … but your “very specific” phrasing makes it sound like there’s a prescription for exactly what is protected by copyright - there is not.

> Ideas are not protected by copyright.

The expression of the idea is, however. Same with facts. The fact that I live at a specific street address is not protected. My sentence construction explaining my specific street address is protected.

> The output of an LLM …

… is not protected, not matter its shape. The US Copyright Office has declared as much.

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#374

Earlier quoted context omitted.

1. Some people want that, with good reason. 2. There's incredible value in what they stole. 3. IANAL, but I don't believe "but now everyone can write like a terrible version of the writer we fleeced" is a valid legal defense.

> stole Copying is not theft.

AFAICT it is if you're not a trillion dollar company.

https://youtu.be/ALZZx1xmAzg?si=ugquA7uKT3ABGdws

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#375

Earlier quoted context omitted.

> Add regulation/enforcement to the big companies and you often shut out the smaller ones following. That is the case, any regulation increases the cost to enter a market. But in this case, its irrelevant because the moat of cost to enter is already unfathomable and secondly, they are not adding regulation but fining them for committing a crime. So yeah, adding that every food compnay needs 3 health inspectors that t…

>the moat of cost to enter is already unfathomable At the moment.

There are multiple ways to respond to that and I will try and summarise them.

Current believe is that its a "winner takes all market", so companies are acting rationally and using Brute Force compute to get there first. Training costs scale linearly, which means the moat is directly related to compute cost

There are theories that they are wasting 90% of training costs and there are more efficient ways to do it than throw compute at the problem. But if thats the case then chances are the market is not "winner takes all". Which then means the valuation of the ENTIRE market is overvalued.

Basically the only way for the assertion "at the moment" to be true is if the market is a bubble, else if the current theory of winner takes all market means a monopoly will make it so that cost isnt even the worst of the moats to enter.

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#377
post #230

Earlier quoted context omitted.

It is legal to train LLMs on books but illegal to train on output of LLMs. Perfect - an absolute steal for 1.5B.

> but illegal to train on output of LLMs. Since when?

Typically the big LLM providers write in the their ToS that it is prohibited to use their output to train another LLM.

Whereas for a book it is fair use.

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#378
post #293

A one time payment like 1.5B doesn’t do anything. There needs to be a royalty payment based on if the AI regurgitates existing ideas. That is probably the correct way to legislate this. If anything a human does can instantly be copied by an LLM, and then sent to all its subscribers, things need to change

This settlement has basically nothing to do with LLMs. At least not as far as the courts are concerned. Alsup ruled [0] that feeding a book into an LLM is transformative and counts as fair use. Especially when they purchased a physical copy of the book, scanned it, and destroyed the original. But if I'm reading the ruling correctly, Anthropic might have been fine even with feeding pirated books into their LLM (as lon…

Who would have thought, that this is the way, which we take to arrive at the burning books stage again? They neatly line up with historical perpetrators in that regard.

Re: Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

#380

Earlier quoted context omitted.

Why is everyone talking about human analogies when LLMs are not humans?

Why is it different other than, "just cause?" No one seems to have actual reasoning to back it up while it feels very similar the other way around, is human brains and neural nets (notwithstanding that they're both called neurons) seem to learn similarly and can act on similar classes of problems like language and mathematics.

are you asking what the diff is between a human and an LLM?
Post reply on HN