Live data from Hacker News

Nvidia contacted Anna's Archive to access books

torrentfreak.com

131–140 of 160 posts

Re: Nvidia contacted Anna's Archive to access books

#131

Earlier quoted context omitted.

Yes, it's been discussed many times before. All the corporations training LLMs have to have done a legal analysis and concluded that it's defensible. Even one of the white papers commissioned by the FSF ( " Copyright Implications of the Use of Code Repositories to Train a Machine Learning Model " at https://www.fsf.org/licensing/copilot/copyright-implications... ), concluded that using copyrighted data to train AI wa…

So it's legal to train an "intelligence" on everything for free based on fair use, but it's not legal to train another intelligence (my brain) on it?

No, it's also not illegal to train your brain. If you break into a store, and read all the books, you'll get arrested for breaking and entering. Not for reading the books. My (superficial) take on the argument is that they're hoping by saying "it's not illegal to read" no one will notice, and no one will ask how they got into the book store to begin with.

Re: Nvidia contacted Anna's Archive to access books

#132
post #90

Earlier quoted context omitted.

Did you pirated this movie? No I did not, it is fair use because this movie is nothing more than a statistical correlation to my dopamine production.

>Did you pirated this movie? No I did not, [...] You're probably being sarcastic but that's actually how the law works. You'll note that when people get sued for "pirating" movies, it's almost always because they were caught seeding a torrent, not for the act of watching an illegal copy. Movie studios don't go after visitors of illegal streaming sites, for instance.

It's how the law works for those at the top of the oligarchy.

Re: Nvidia contacted Anna's Archive to access books

#133

Earlier quoted context omitted.

Okay, so go check out 500 TB worth of books from the library. I'll wait

If I’m rich enough to employ thousands of people I can hire each one of them to borrow as many books as possible then use all the books to train an AI. Perfectly legal. And also very possible. Point being that the library prevents you from checking out 500gb because of logistical issues. First how can you carry all those books and how can they let other patrons in the library check out books if you grabbed that many?…

Great! Then it's perfectly legal.

As long as you obtain the books legally then it's legal

This really isn't that hard

Re: Nvidia contacted Anna's Archive to access books

#134

Earlier quoted context omitted.

If I’m rich enough to employ thousands of people I can hire each one of them to borrow as many books as possible then use all the books to train an AI. Perfectly legal. And also very possible. Point being that the library prevents you from checking out 500gb because of logistical issues. First how can you carry all those books and how can they let other patrons in the library check out books if you grabbed that many?…

Great! Then it's perfectly legal. As long as you obtain the books legally then it's legal This really isn't that hard

So you’re wrong when you said you have to pay for the books. You don’t.

Re: Nvidia contacted Anna's Archive to access books

#135
post #31

Earlier quoted context omitted.

What do you mean 'sucked up'? It's data on their machines already, people willingly give them the data, so Amazon can process and offer it to readers. No sucking needed, just use the data people uploaded to you already.

There's definitely a legal & contractual difference between (1) storing the books on your servers in order to provide them to end users who have purchased licenses to read them and (2) using that same data for training a model that might be used to create books that compete with the originals. I'm pretty sure that's why GP means by "sucking up." This is analogous the difference between Gmail using search within your…

They may not serve ads but you don't know they don't train their models on them.

If I still used Gmail I'd read the terms of service real close.

Re: Nvidia contacted Anna's Archive to access books

#136

Earlier quoted context omitted.

But to train the models they have to download it first (make a copy)

You had to do this for reading too. The words were burned onto your retina as volatile memory before getting processed by your brain. You retina likely overwrote it's "memory" as soon as you looked at something else, but that's no different than copying and deleting or the more apt analogy: streaming.

BS. Nvidia store use the copy for each training run, or do you really thing the just download it each time in real time for training?

Re: Nvidia contacted Anna's Archive to access books

#137

Earlier quoted context omitted.

Did you pirated this movie? No I did not, it is fair use because this movie is nothing more than a statistical correlation to my dopamine production.

Note that what copyright law prohibits is the action of producing a copy for someone else, not the action of obtaining a copy for yourself.

If I am not mistaken, the law prohibits producing any unauthorized copies. So if you download a pirated book on a computer, you produce an illegal copy: [1]. If I am not missing anything, ML companies are galaxy-scale infringers.

> 106. Exclusive rights in copyrighted works

> Subject to sections 107 through 122, the owner of copyright under this title has the exclusive rights to do and to authorize any of the following:

> (1) to reproduce the copyrighted work in copies or phonorecords;

> 501. Infringement of copyright

> (a) Anyone who violates any of the exclusive rights of the copyright owner as provided by sections 106 through 122 or of the author as provided in section 106A(a), or who imports copies or phonorecords into the United States in violation of section 602, is an infringer of the copyright or right of the author, as the case may be.

[1] https://www.copyright.gov/title17/92chap5.html

Re: Nvidia contacted Anna's Archive to access books

#138

Earlier quoted context omitted.

Yes, it's been discussed many times before. All the corporations training LLMs have to have done a legal analysis and concluded that it's defensible. Even one of the white papers commissioned by the FSF ( " Copyright Implications of the Use of Code Repositories to Train a Machine Learning Model " at https://www.fsf.org/licensing/copilot/copyright-implications... ), concluded that using copyrighted data to train AI wa…

So it's legal to train an "intelligence" on everything for free based on fair use, but it's not legal to train another intelligence (my brain) on it?

You're close to an important point.

Our current laws are written to make it legal for you to copy the Quran via your brain — some people learn it by rote and can stand up and speak the entire work from one end to the other. This is intended to be legal. Fair use of the Quran.

I went to a concert recently where someone copied every word and (as far as I could hear) every note from a copyrighted work by Bruce Springsteen. Singing and playing. This too is intended to be fair use.

You can learn how to play and sing Springsteen songs verbatim, and you can use his records to learn to sound like him when you sing, and that's intended to be legal.

Since the law doesn't say "but you cannot write a program to do these things, or run such a program once written", why would it be illegal to do the same thing using some code?

The people who want the law to differentiate have a difficult challenge in front of them. As I see it, they need to differentiate between what humans do to learn from what machines do, and that implies really knowing what humans do. And then they need to draw boundaries, making various kinds of computer-assisted human learning either legal or illegal.

Some of them say things like "when an AI draws Calvin and Hobbes in the style of Breughel, it obviously has copied paintings by Breughel" but a court will ask why that's obvious. Is it really obvious that the way it does that drawing necessarily involves copying, when you as a human can do the same thing without copying?

Re: Nvidia contacted Anna's Archive to access books

#139

Earlier quoted context omitted.

But to train the models they have to download it first (make a copy)

You had to do this for reading too. The words were burned onto your retina as volatile memory before getting processed by your brain. You retina likely overwrote it's "memory" as soon as you looked at something else, but that's no different than copying and deleting or the more apt analogy: streaming.

The law makes a distinction between storing it on a disk and just remembering the content. The latter is not a "copy" and not a subject of law:

> “Copies” are material objects, other than phonorecords, in which a work is fixed by any method now known or later developed, and from which the work can be perceived, reproduced, or otherwise communicated, either directly or with the aid of a machine or device. The term “copies” includes the material object, other than a phonorecord, in which the work is first fixed.

> A work is “fixed” in a tangible medium of expression when its embodiment in a copy or phonorecord, by or under the authority of the author, is sufficiently permanent or stable to permit it to be perceived, reproduced, or otherwise communicated for a period of more than transitory duration. A work consisting of sounds, images, or both, that are being transmitted, is “fixed” for purposes of this title if a fixation of the work is being made simultaneously with its transmission.

https://www.copyright.gov/title17/92chap1.html

Post reply on HN