Live data from Hacker News

Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

businessinsider.com

421–430 of 686 posts

Re: Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

#421

Earlier quoted context omitted.

you've created a very obvious category mistake in your final summary by confusing intellectual property--which can be copied at no penalty to an owner (except nebulous 'alternate universe' theories)--with actual property, and a farmer and his land, with a crop that cannot be enjoyed twice. you're saying copying a book is worse than robbing a farmer of his food and/or livelihood, which cannot be replaced to duplicated…

With this special infinite-land-land though, what's special about the farmer's land is that he's expended energy to make it that way, just as the author has expended energy to find his text. Just as the farmer obtains his livelihood from the investment-of-energy-to-raise-crops-to-energy cycle the author has his livelihood by the investment-of-energy-to-finding-a-useful-work-to-energy cycle. So he is in fact robbed in…

You're saying that a copy of a digital thing is the same as the "only" of a physical thing. But that's not true. You can't sell grain twice, but you can sell a movie many times (especially when you account for format changes, remasterings, platform locks, licensing for special usecases like remixing, broadcasts, etc).

You'd have to steal the author's ownership of the intellectual property in order for the comparison to be valid, just as you stole ownership of his crop.

Separately, there is a reason why theft and copyright infringement are two distinct concepts in law.

Re: Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

#422
post #385

Earlier quoted context omitted.

You make some good points, this really is going to take some careful judgment and chances are it's too complex for an actual courtroom to yield an ideal outcome. Now places like Flea markets have been known to have a counterfeit DVD or two. And there is more than one way to compare to non-digital content. Regular books and periodicals can be sold out and/or out-of-print, but digital versions do not have these same ex…

> If you save it from the bin you should be able to do whatever you want with it either way, scan it how you see fit. Don’t we already have laws covering this? For example, sometimes excess books can be thrown in the bin. Often, they have the covers removed. Some will say something to the effect that “if you’ve received this without a cover it is a copyright violation.” I think one of the points of the lawsuit is it…

Another good point, and it's a fine point as well.

You could split hairs over whether saving an item from the bin occurred after a procedure to remove covers and it was already dumped, or before any contemplation was made about if or when dumping would take place.

Saving either way would be preserving what would otherwise be lost, even if it was well premeditated in advance of any imminent risk.

What if it was the last remaining copy?

Or even the only copy ever in existence of an original manuscript?

It's just not a concept suitable for a black & white judgment.

That's a very good sign that probably an entire book of regulations needs to be thrown out instead, and a new law written to replace it with something more sensible.

Re: Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

#423
post #253

Earlier quoted context omitted.

That a service incorporating the authors' works exists is not at issue. The plaintiffs' claims are, as summarized by Alsup: First, Authors argue that using works to train Claude’s underlying LLMs was like using works to train any person to read and write, so Authors should be able to exclude Anthropic from this use (Opp. 16). Second, to that last point, Authors further argue that the training was intended to memorize…

The first paragraph sounds absurd, so I looked into the PDF, and here's the full version I found: > First, Authors argue that using works to train Claude’s underlying LLMs was like using works to train any person to read and write, so Authors should be able to exclude Anthropic from this use (Opp. 16). But Authors cannot rightly exclude anyone from using their works for training or learning as such. Everyone reads te…

For everyone arguing that there’s no harm in anthropomorphizing an LLM, witness this rationalization. They talk about training and learning as if this is somehow comparable to human activities. The idea that LLM training is comparable to a person learning seems way out there to me.

“We have admired, memorized, and internalized their sweeping themes, their substantive points, and their stylistic solutions to recurring writing problems.”

Claude is not doing any of these things. There is no admiration, no internalizing of sweeping themes. There’s a network encoding data.

We’re talking about a machine that accepts content and then produces more content. It’s not a person, it’s owned by a corporation that earns money on literally every word this machine produces. If it didn’t have this large corpus of input data (copyrighted works) it could not produce the output data for which people are willing to pay money. This all happens at a scale no individual could achieve because, as we know, it is a machine.

Re: Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

#424
post #414

Earlier quoted context omitted.

> That is, he ruled that > - buying, physically cutting up, physically digitizing books, and using them for training is fair use > - pirating the books for their digital library is not fair use. That seems inconsistent with one another. If it's fair use, how is it piracy? It also seems pragmatically trash. It doesn't do the authors any good for the AI company to buy one copy of their book (and a used one at that), bu…

These are two separate actions that Anthropic did: * They downloaded a massive online library of pirated books that someone else was distributing illegally. This was not fair use. * They then digitised a bunch of books that they physically owned copies of. This was fair use. This part of the ruling is pretty much existing law. If you have a physical book (or own a digital copy of a book), you can largely do what you…

> This part of the ruling is pretty much existing law. If you have a physical book (or own a digital copy of a book), you can largely do what you like with it within the confines of your own home, including digitising it. But you are not allowed to distribute those digital copies to others, nor are you allowed to download other people's digital copies that you don't own the rights to.

Can you point me to the US Supreme Court case where this is existing law?

It's pretty clear that if you have a physical copy of a book, you can lend it to someone. It also seems pretty reasonable that the person borrowing it could make fair use of it, e.g. if you borrow a book from the library to write a book review and then quote an excerpt from it. So the only thing that's left is, what if you do the same thing over the internet?

Shouldn't we be able to distinguish this from the case where someone is distributing multiple copies of a work without authorization and the recipients are each making and keeping permanent copies of it?

Re: Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

#425

Earlier quoted context omitted.

Computers cannot learn and are not subjects to laws. What happens, is a human takes a copyrighted work, makes an unauthorized digital copy, and loads it into a computer without authorization from copyright owner.

And they are not selling this or distributing this. The model is very different.

I have to disagree, without all the copyrighted input data there would be no output data for these companies to sell. This output data is the product and they are distributing it for dollars.

Re: Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

#426
The solution has always been: show us the training data.

As a researcher I've been furious that we publish papers where the research data is unknown. To add insult to injury, we have the audacity to start making claims about "zero-shot", "low-shot", "OOD", and other such things. It is utterly laughable. These would be tough claims to make *even if we knew the data*, simply because of its size. But not knowing the data, it is outlandish. Especially because the presumptions are "everything on the internet." It would be like training on all of GitHub and then writing your own simple programming questions to test an LLM[0]. Analyzing that amount of data is just intractable, and we currently do not have the mathematical tools to do so. But this is a much harder problem to crack when we're just conjecturing and ultimately this makes interoperability more difficult.

On top of all of that, we've been playing this weird legal game. Where it seems that every company has had to cheat. I can understand how smaller companies turn to torrenting to compete, but when it is big names like Meta, Google, Nvidia, OpenAI (Microsoft), etc it is just wild. This isn't even following the highly controversial advice of Eric Schmidt "Steal everything, then if you get big, let the lawyers figure it out." This is just "steal everything, even if you could pay for it." We're talking about the richest companies in the entire world. Some of the, if not the, richest companies to ever exist.

Look, can't we just try to be a little ethical? There is, in fact, enough money to go around. We've seen unprecedented growth in the last few years. It was only 2018 when Apple became the first trillion dollar company, 2020 when it became the second two trillion, and 2022 when it became the first three trillion dollar company. Now we have 10 companies north of the trillion dollar mark![3] (5 above $2T and 3 above $3T) These values have exploded in the last 5 years! It feels difficult to say that we don't have enough money to do things better. To at least not completely screw over "the little guy." I am unconvinced that these companies would be hindered if they had to broker some deal for training data. Hell, they're already going to war over data access.

My point here is that these two things align. We're talking about how this technology is so dangerous (every single one of those CEOs has made that statement) and yet we can't remain remotely ethical? How can you shout "ONLY I CAN MAKE SAFE AI" while acting so unethically? There's always moral gray areas but is this really one of them? I even say this as someone who has torrented books myself![4] We are holding back the data needed to make AI safe and interpretable while handing the keys to those who actively demonstrate that they should not hold the power. I don't understand why this is even that controversial.

[0] Yes, this is a snipe at HumanEval. Yes, I will make the strong claim that the dataset was spoiled from day 1. If you doubt it, go read the paper and look at the questions (HuggingFace).

[1] https://www.theverge.com/2024/8/14/24220658/google-eric-schm...

[2] https://en.wikipedia.org/wiki/List_of_public_corporations_by...

[3] https://companiesmarketcap.com/

[4] I can agree it is wrong, but can we agree there is a big difference between a student torrenting a book and a billion/trillion dollar company torrenting millions of books? I even lean on the side of free access to information, and am a fan of Aaron Swartz and SciHub. I make all my works available on ArXiv. But we can recognize there's a big difference between a singular person doing this at a small scale and a huge multi-national conglomerate doing it at a large scale. I can't even believe we so frequently compare these actions!

Re: Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

#427

Earlier quoted context omitted.

As they mentioned, the piracy part is obvious. It's the fair use part that will set an important precedent for being able to train on copyrighted works as long as you have legally acquired a copy.

Cue physical books being licensed not sold in the futur with restricted agreements …

Also music, videos, photos, etc.

Re: Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

#428
post #396

Earlier quoted context omitted.

That's right, so I can't individually discuss terms with each and every media creator, so from now on, I can just pirate everything.

Needing a copy of one book you're going to spend a week reading has a lot less overhead than needing a copy of every book that you're going to process with a computer in bulk.

I like to glance at the cover art. I can do ten per second when I really get into my flow state. Sometimes I read them also, but that's incidental.

Re: Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

#429
post #380

Earlier quoted context omitted.

> buying, physically cutting up, physically digitizing books, and using them for training is fair use So Suno would only really need to buy the physical albums and rip them to be able to generate music at an industrial scale?

Only if the physical albums don't have copy protection, otherwise you're circumenventing it and that's illegal. Or is it, against the right to private copy? If anything, AI at least shows that all of the existing copyright laws are utter bullshit made to make Disney happy. Do keep in mind though: this is only for the wealthy. They're still going to send the Pinkertons at your house if you dare copy a Blu-ray.

> They're still going to send the Pinkertons at your house if you dare copy a Blu-ray.

Hey woah now, that's a Hasbro play, not a Disney one.

Re: Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

#430

Apparently it's a common business practice. Spotify (even though I can't find any proof) seems to have build their software and business on pirated music. There is some more in this Article [0]. https://torrentfreak.com/spotifys-beta-used-pirate-mp3-files... Funky quote: > Rumors that early versions of Spotify used ‘pirate’ MP3s have been floating around the Internet for years. People who had access to the service in…

Google Music originally let people upload their own digital music files. The argument at the time was that whether or not the files were legally obtained was not Google’s problem. I believe Amazon had a similar service.

https://www.computerworld.com/article/1447323/google-reporte...

Post reply on HN