Live data from Hacker News

Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

businessinsider.com

451–460 of 686 posts

Re: Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

#451

Earlier quoted context omitted.

> Simply, if the models can think then it is no different than a person reading many books and building something new from their learnings. No, that's fallacious. Using anthropomorphic words to describe a machine does not give it the same kinds of rights and affordances we give real people.

Actually, it does, at least for this case. The judge just said so.

[deleted]

Re: Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

#452

Earlier quoted context omitted.

> Downloading a document is fine as long as it is not shared outside? I've fixed your question so that it accurately represents what I said and doesn't put words in my mouth. If I click on a link and download a document, is that illegal? I do not know if the person has the right to distribute it or not. IANAL, but when people were getting sued by the RIAA years back, it was never about downloading, but also distribut…

> it was never about downloading, but also distribution. Did you mean to write "but about distribution" here?

Yes, thank you for catching that. Unfortunately, I cannot edit it now.

Re: Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

#453
post #380

Earlier quoted context omitted.

You skipped quotes about the other important side: > But Alsup drew a firm line when it came to piracy. > "Anthropic had no entitlement to use pirated copies for its central library," Alsup wrote. "Creating a permanent, general-purpose library was not itself a fair use excusing Anthropic's piracy." That is, he ruled that - buying, physically cutting up, physically digitizing books, and using them for training is fair…

> buying, physically cutting up, physically digitizing books, and using them for training is fair use So Suno would only really need to buy the physical albums and rip them to be able to generate music at an industrial scale?

If it's fair use to train a model, that doesn't necessarily imply that the model can be legally used to generate anything.

Re: Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

#454

actual title: "Anthropic cut up millions of used books to train Claude — and downloaded over 7 million pirated ones too, a judge said." A not-so-subtle difference. That said, in a sane world, they shouldn't have needed to cut up all those used books yet again when there's obviously already an existing file that does all the work.

The importance of acquiring the physical book was the transfer of compensation to the author.

You're not wrong, but that's one heck of a way to do it. It involves the destruction of 7 million books, which ... I really don't quite see the "promotion of Progress of Science and useful Arts" in that.

Re: Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

#456

Earlier quoted context omitted.

It kind of is though? It's not the only reason fair use exists, but it's the thing that allows e.g. search engines to exist, and that seems pretty important. > And you ever heard of a publishing house? They don't need to negotiate with every single author individually. That's preposterous. There are thousands of publishing houses and millions of self-published authors on top of that. Many books are also out of print…

>It kind of is though? No, it kinda isn't. Show me anything that supports this idea beyond your own immediate conjecture right now. >It's not the only reason fair use exists, but it's the thing that allows e.g. search engines to exist, and that seems pretty important. No, that's the transformative element of what a search engine provides. Search engines are not legal because they can't contact each licensor, they are…

> Show me anything that supports this idea beyond your own immediate conjecture right now

It's inherent in the nature of the test. The most important fair use factor is the effect on the market for the work, so if the use would be uneconomical without fair use then the effect on the market is negligible because the alternative would be that the use doesn't happen rather than that the author gets paid for it.

> No, that's the transformative element of what a search engine provides. Search engines are not legal because they can't contact each licensor, they are legal because they are considered hugely transformative features.

To make a search engine you have to do two things. One is to download a copy of the whole internet, the other is to create a search index. I'm talking about the first one, you're talking about the second one.

> Okay, and? How many customers does Microsoft bill on a monthly basis?

Microsoft does this with an automated system. There is no single automated system where you can get every book ever written, and separately interfacing with all of the many systems needed in order to do it is the source of the overhead.

Re: Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

#457
post #414

Earlier quoted context omitted.

These are two separate actions that Anthropic did: * They downloaded a massive online library of pirated books that someone else was distributing illegally. This was not fair use. * They then digitised a bunch of books that they physically owned copies of. This was fair use. This part of the ruling is pretty much existing law. If you have a physical book (or own a digital copy of a book), you can largely do what you…

> This part of the ruling is pretty much existing law. If you have a physical book (or own a digital copy of a book), you can largely do what you like with it within the confines of your own home, including digitising it. But you are not allowed to distribute those digital copies to others, nor are you allowed to download other people's digital copies that you don't own the rights to. Can you point me to the US Supre…

It is “established” law because the Copyright Act itself and a string of unanimous or near-unanimous appellate decisions (google ReDigi on digital transfers and Sony and the first-sale for personal use and physical lending) uniformly apply the same principles, leaving no circuit split and no conflicting precedent for the Supreme Court to resolve. In the U.S. system statutory text interpreted consistently by the Courts of Appeals becomes binding law nationwide unless and until the Supreme Court or Congress says otherwise.

Re: Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

#458
post #374

Earlier quoted context omitted.

But the AI used the content to learn how to copy and recreate it. Is ‘re-creation’ a better concept for us? People already use pirated software for product creation. Hypothetical: I know a guy who learned photoshop on a pirated copy of Photoshop. He went on to be a graphic designer. All his earnings are ‘proceeds from crime’ He never used the pirated software to produce content.

So can we officially download pirated content to learn stuff now?

How often does a link get posted here of content that is behind a paywall? If you bypass it to read it, didny't you just learn via illegal content? I'm not sure where the "official" comes in, but it's clearly widely accepted.

If you watch a YouTube video to learn something and it's later taken down for using copyrighted images, you learned from illegal content.

Re: Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

#459

Earlier quoted context omitted.

From my understanding: > pirating the books for their digital library is not fair use. "Pirating" is a fuzzy word and has no real meaning. Specifically, I think this is the cruz: > without adding new copies, creating new works, or redistributing existing copies Essentially: downloading is fine, sharing/uploading up is not. Which makes sense. The assertion here is that Anthropic (from this line) did not distribute the…

The legal context here is that "format shifting" has not previously been held to be sufficient for fair use on its own, and downloading for personal use has also been considered infringing. Just look at the numerous media industry lawsuits against individuals that only mention downloading, not sharing for examples. It's a bit surprising that you can suddenly download copyrighted materials for personal use and and it'…

> the numerous media industry lawsuits against individuals that only mention downloading,

I never saw any of these. All the cases I saw were related to people using torrents or other P2P software (which aren't just downloading). These might exist, but I haven't seen them.

> It's a bit surprising that you can suddenly download copyrighted materials for personal use and it's kosher as long as you don't share them with others.

Every click on a link is a risk of downloading copyrighted material you don't have the rights to.

Searching the internet, it appears that it's a civil infraction, but it's also confused with the notion that "piracy" is illegal, a term that's used for many different purposes. I see "It is illegal to download any music or movies that are copyrighted." under legal advice, which I know as a statement is not true.

Hence my confusion.

I should note: I'm not arguing from the perspective of whether it's morally or ethically right. Only that even in the context of this thread, things are phrased that aren't clear.

Re: Anthropic cut up millions of used books, and downloaded 7M pirated ones – judge

#460

Earlier quoted context omitted.

The first paragraph sounds absurd, so I looked into the PDF, and here's the full version I found: > First, Authors argue that using works to train Claude’s underlying LLMs was like using works to train any person to read and write, so Authors should be able to exclude Anthropic from this use (Opp. 16). But Authors cannot rightly exclude anyone from using their works for training or learning as such. Everyone reads te…

For everyone arguing that there’s no harm in anthropomorphizing an LLM, witness this rationalization. They talk about training and learning as if this is somehow comparable to human activities. The idea that LLM training is comparable to a person learning seems way out there to me. “We have admired, memorized, and internalized their sweeping themes, their substantive points, and their stylistic solutions to recurring…

There may be no admiration, but there definitely is an internalising of sweeping themes, and all the other things in your quotation, which anyone can fetch by asking it for the themes/substantive points/stylistic solutions of one of the books it has (for lack of a better verb) read.

That the mechanism performing these things is a network encoding data is… well, that description, at that level of abstraction, is a similarity with the way a human does it, not even a difference.

My network is a 3D mess made of pointy bi-lipid bags exchanging protons across gaps moderated by the presence of neurochemicals, rather than flat sheets of silicon exchanging electrons across tuned energy band-gaps moderated by other electrons, but it's still a network.

> We’re talking about a machine that accepts content and then produces more content. It’s not a person, it’s owned by a corporation that earns money on literally every word this machine produces. If it didn’t have this large corpus of input data (copyrighted works) it could not produce the output data for which people are willing to pay money. This all happens at a scale no individual could achieve because, as we know, it is a machine.

My brain is a machine that accepts content in the form of job offers and JIRA tickets (amongst other things), and then produces more content in the form of pull requests (amongst other things). For the sake specifically of this question, do the other things make a difference? While I count as a person and am not owned by any corporation, when I work for one, they do earn money on the words this biological machine produces. (And given all the models which are free to use, the LLMs definitely don't earn money on "literally" every word those models produce). If I didn't have the large corpus of input data — and there absolutely was copyright on a lot of the school textbooks and the TV broadcast educational content of the 80s and 90s when I was at school, and the Java programming language that formed the backbone of my university degree — I could not produce the output data for which people are willing to pay money.

Should corporations who hire me be required to pay Oracle every time I remember and use a solution that I learned from a Java course, even when I'm not writing Java?

That the LLMs do this at a scale no individual could achieve because it is a machine, means it's got the potential to wipe me out economically. Economics threat of automation has been a real issue at least since the luddites if not earlier, and I don't know how the dice will fall this time around, so even though I have one layer of backup plan, I am well aware it may not work, and if it doesn't then government action will have to happen because a lot of other people will be in trouble before trouble gets to me (and recent history shows that this doesn't mean "there won't be trouble").

Copyright law is one example of government action. So is mandatory education. So is UBI, but so too is feudalism.

Good luck to us all.

Post reply on HN