Live data from Hacker News

Nvidia contacted Anna's Archive to access books

torrentfreak.com

41–50 of 160 posts

Re: Nvidia contacted Anna's Archive to access books

#41

Earlier quoted context omitted.

> To call training illegal is similar to calling reading a book and remembering it illegal. Perhaps, but reproducing the book from this memory could very well be illegal. And these models are all about production.

Models don’t reproduce books though. It’s impossible for a model to reproduce something word for word because the model never copied the book. Most of the best fit curve runs along a path that doesn’t even touch an actual data point.

If there is one exact sentence taken out of the book and not referenced in quotes and exact source, that triggers copyright laws. So model doesnt have to reproduce the entire book, it only required to reproduce one specific sentence (which may be a characteristic sentence to that author or to that book).

Re: Nvidia contacted Anna's Archive to access books

#42
post #10

It's generous of them to ask for permission.

They wanted access to a faster pipe to slurp 500 terabytes, and that access comes at a cost. It wasn’t about permission. And yeah they should be sued into the next century for copyright infringement. $4Trillion company illegally downloading the entire corpus of published literature for reuse is clearly infringement, its an absurdity to say that it’s fair use just to look for statistical correlations when training LLM…

Whatever they get sued for would be pocket change.

Re: Nvidia contacted Anna's Archive to access books

#43

Earlier quoted context omitted.

It does make sense. It’s controversial. Your memory memorizes things in the same way. So what nvidia does here is no different, the AI doesn’t actually copy any of the books. To call training illegal is similar to calling reading a book and remembering it illegal. Our copyright laws are nowhere near detailed enough to specify anything in detail here so there is indeed a logical and technical inconsistency here. I can…

> To call training illegal is similar to calling reading a book and remembering it illegal. A type of wishful thinking fallacy. In law scale matters. It's legal for you to possess a single joint. It's not legal to possess 400 tons of weed in a warehouse.

It is not the scale that matters here, in your example, but intent. With 1 joint, you want to smoke yourself. With 400, you very possibly want to sell it to others. Scale in itself doesnt matter, scale matters only as to the extent it changes what your intention may be.

Re: Nvidia contacted Anna's Archive to access books

#44

Earlier quoted context omitted.

Models don’t reproduce books though. It’s impossible for a model to reproduce something word for word because the model never copied the book. Most of the best fit curve runs along a path that doesn’t even touch an actual data point.

They do memorize some books. You can test this trivially by asking ChatGPT to produce the first chapter of something in the public domain -- for example a Tale of Two Cities. It may not be word for word exact, but it'll be very close. These academics were able to get multiple LLMs to produce large amounts of text from Harry Potter: https://arxiv.org/abs/2601.02671

In that case I would say it is the act of reproducing the books that is illegal. Training the AI on said books is not.

So the illegality rests at the point of output and not at the point of input.

I’m just speaking in terms of the technical interpretation of what’s in place. My personal views on what it should be are another topic.

Re: Nvidia contacted Anna's Archive to access books

#45
post #3

> In response, NVIDIA defended its actions as fair use, noting that books are nothing more than statistical correlations to its AI models. Does this even make sense? Are the copyright laws so bad that a statement like this would actually be in NVIDIA’s favor?

> Does this even make sense? Are the copyright laws so bad that a statement like this would actually be in NVIDIA’s favor? It makes some sense, yeah. There's also precedent, in google scanning massive amounts of books, but not reproducing them. Most of our current copyright laws deal with reproductions. That's a no-no. It gets murky on the rest. Nvda's argument here is that they're not reproducing the works, they're…

Scanning books is literally reproducing them. Copying books from Anna's Archive is also literally reproducing them. The idea that it is only copyright infringement if you engage in further reproduction is just wrong.

As a consumer you are unlikely to be targeted for such "end-user" infringement, but that doesn't mean it's not infringement.

Re: Nvidia contacted Anna's Archive to access books

#46

Earlier quoted context omitted.

It does make sense. It’s controversial. Your memory memorizes things in the same way. So what nvidia does here is no different, the AI doesn’t actually copy any of the books. To call training illegal is similar to calling reading a book and remembering it illegal. Our copyright laws are nowhere near detailed enough to specify anything in detail here so there is indeed a logical and technical inconsistency here. I can…

You can only read the book, if you purchased it. Even if you dont have the intent to reproduce it, you must purchase it. So, I guess NVDA should just purchase all those books, no?

Yep, I agree. That’s the part that’s clearly illegal. They should purchase the books, but they didn’t.

Re: Nvidia contacted Anna's Archive to access books

#47

Earlier quoted context omitted.

> To call training illegal is similar to calling reading a book and remembering it illegal. A type of wishful thinking fallacy. In law scale matters. It's legal for you to possess a single joint. It's not legal to possess 400 tons of weed in a warehouse.

It is not the scale that matters here, in your example, but intent. With 1 joint, you want to smoke yourself. With 400, you very possibly want to sell it to others. Scale in itself doesnt matter, scale matters only as to the extent it changes what your intention may be.

It’s clear nvidia and every single one of these big AI corps do not want their AIs to violate the law. The intent is clear as day here.

Scale is only used for emergence, openAI found that training transformers on the entire internet would make is more then just a next token predictor and that is the intent everyone is going for when building these things.

Re: Nvidia contacted Anna's Archive to access books

#48
post #3

> In response, NVIDIA defended its actions as fair use, noting that books are nothing more than statistical correlations to its AI models. Does this even make sense? Are the copyright laws so bad that a statement like this would actually be in NVIDIA’s favor?

> Does this even make sense? Are the copyright laws so bad that a statement like this would actually be in NVIDIA’s favor? It makes some sense, yeah. There's also precedent, in google scanning massive amounts of books, but not reproducing them. Most of our current copyright laws deal with reproductions. That's a no-no. It gets murky on the rest. Nvda's argument here is that they're not reproducing the works, they're…

> I don't see how they get around "procuring them" from 3rd party dubious sources

Yeah, isn't this what Anthropic was found guilty off?

Re: Nvidia contacted Anna's Archive to access books

#49
post #3

> In response, NVIDIA defended its actions as fair use, noting that books are nothing more than statistical correlations to its AI models. Does this even make sense? Are the copyright laws so bad that a statement like this would actually be in NVIDIA’s favor?

It does make sense. It’s controversial. Your memory memorizes things in the same way. So what nvidia does here is no different, the AI doesn’t actually copy any of the books. To call training illegal is similar to calling reading a book and remembering it illegal. Our copyright laws are nowhere near detailed enough to specify anything in detail here so there is indeed a logical and technical inconsistency here. I can…

But it’s not just about recall and reproduction. If they used Anna’s Archive the books were obtained and copied without a license, before they were fed in as training data.

Re: Nvidia contacted Anna's Archive to access books

#50

Earlier quoted context omitted.

They do memorize some books. You can test this trivially by asking ChatGPT to produce the first chapter of something in the public domain -- for example a Tale of Two Cities. It may not be word for word exact, but it'll be very close. These academics were able to get multiple LLMs to produce large amounts of text from Harry Potter: https://arxiv.org/abs/2601.02671

In that case I would say it is the act of reproducing the books that is illegal. Training the AI on said books is not. So the illegality rests at the point of output and not at the point of input. I’m just speaking in terms of the technical interpretation of what’s in place. My personal views on what it should be are another topic.

> So the illegality rests at the point of output and not at the point of input.

It's not as simple as that, as this settlement shows [1].

Also, generating output is what these models are primarily trained for.

[1]: https://www.bbc.com/news/articles/c5y4jpg922qo

Post reply on HN