Live data from Hacker News

Nvidia contacted Anna's Archive to access books

torrentfreak.com

81–90 of 160 posts

Re: Nvidia contacted Anna's Archive to access books

#81
post #15

Sounds like BS. Why would nvidia need the books. Do they even have a chatbot? I doubt the books help with framegen.

From the top of the linked article:

    > NVIDIA is also developing its own models, including NeMo, Retro-48B, InstructRetro, and Megatron. These are trained using their own hardware and with help from large text libraries, much like other tech giants do.
You can download the models here: https://huggingface.co/nvidia

Re: Nvidia contacted Anna's Archive to access books

#82

Earlier quoted context omitted.

> To call training illegal is similar to calling reading a book and remembering it illegal. A type of wishful thinking fallacy. In law scale matters. It's legal for you to possess a single joint. It's not legal to possess 400 tons of weed in a warehouse.

Er no. I’ve read and remember hundreds of books in my life time. It’s not any more illegal based off scale. The law doesn’t differentiate whether I remember one book or a hundred then there’s no difference for thousands or millions. No wishful thinking here.

> Er no. I’ve read and remember hundreds of books in my life time. It’s not any more illegal based off scale.

I'm not sure you understood what you said, but superficially it appears that you are agreeing with me?

Just because it's legal to read 100s of books does not make it legal to slurp up every single piece of produced content ever recorded.

We're talking man many orders of magnitude in scale there, and you're the one who pointed out that scale :-/

Re: Nvidia contacted Anna's Archive to access books

#83

Earlier quoted context omitted.

> To call training illegal is similar to calling reading a book and remembering it illegal. A type of wishful thinking fallacy. In law scale matters. It's legal for you to possess a single joint. It's not legal to possess 400 tons of weed in a warehouse.

It is not the scale that matters here, in your example, but intent. With 1 joint, you want to smoke yourself. With 400, you very possibly want to sell it to others. Scale in itself doesnt matter, scale matters only as to the extent it changes what your intention may be.

> It is not the scale that matters here, in your example, but intent. With 1 joint, you want to smoke yourself. With 400, you very possibly want to sell it to others. Scale in itself doesnt matter, scale matters only as to the extent it changes what your intention may be.

It sounds then like you're saying that scale does indeed matter in this context, as using every single piece of writing in existence isn't being slurped up purely to learn, it's being slurped up to make a profit.

Do you think they'd be able to offer a usefull LLM if the model was trained only what what an average person could read in a lifetime?

Re: Nvidia contacted Anna's Archive to access books

#84
post #3

> In response, NVIDIA defended its actions as fair use, noting that books are nothing more than statistical correlations to its AI models. Does this even make sense? Are the copyright laws so bad that a statement like this would actually be in NVIDIA’s favor?

It's not settled law as it pertains to LLMs, but, yes, creating a "statistical summary" of a book (consider, e.g., a concordance of Joyce's "Ulysses") is generally protected as fair use. However, illegally accessing pirated books to create that concordance is still illegal.

Re: Nvidia contacted Anna's Archive to access books

#86
I've always wondered about some of the torrent whales with multiple petabytes on private trackers. A lot of the whales auto dl every single new torrent that's uploaded. Perhaps even the sites themselves are allowed to operate as a way to get users to crowd source media.

Re: Nvidia contacted Anna's Archive to access books

#87

Earlier quoted context omitted.

Private reproductions are allowed (e.g. backups). Distributing them non-privately is not.

Backups are permitted (and not for all media) when you legally acquired the source. Scanning a physical book is not a permitted backup, and neither is downloading a book from Anna's archive.

> Scanning a physical book is not a permitted backup

On what basis do you claim that?

You're also missing critical legal context. When a would be consumer downloads pirated media in lieu of purchasing it he damages the would be seller. When my automated web scraper inadvertently archives some pirated content on my local disk no one is financially harmed.

The question is where the boundary between those things lies.

Re: Nvidia contacted Anna's Archive to access books

#88

Just to clarify, the most valuable company in the world refuses to pay for digital media?

I see this sentiment posted quite a bit, but have the publishers made any products available that would allow AI training on their works for payment? A naive approach would be to go to an online bookstore and pay $15 for every book, but then you have copyrighted content that is encrypted, that it's a violation of the DMCA to decrypt. I assume you're expecting that they'll reach out and cut a deal with each publishing…

The product i want doesnt exist too. But if I pirate, straight to Alcataraz I go.

Re: Nvidia contacted Anna's Archive to access books

#89

Earlier quoted context omitted.

Models don’t reproduce books though. It’s impossible for a model to reproduce something word for word because the model never copied the book. Most of the best fit curve runs along a path that doesn’t even touch an actual data point.

If there is one exact sentence taken out of the book and not referenced in quotes and exact source, that triggers copyright laws. So model doesnt have to reproduce the entire book, it only required to reproduce one specific sentence (which may be a characteristic sentence to that author or to that book).

Sure, but that use would easily pass a fair use test, at least in the US.

Re: Nvidia contacted Anna's Archive to access books

#90
post #3

> In response, NVIDIA defended its actions as fair use, noting that books are nothing more than statistical correlations to its AI models. Does this even make sense? Are the copyright laws so bad that a statement like this would actually be in NVIDIA’s favor?

Did you pirated this movie? No I did not, it is fair use because this movie is nothing more than a statistical correlation to my dopamine production.

>Did you pirated this movie? No I did not, [...]

You're probably being sarcastic but that's actually how the law works. You'll note that when people get sued for "pirating" movies, it's almost always because they were caught seeding a torrent, not for the act of watching an illegal copy. Movie studios don't go after visitors of illegal streaming sites, for instance.

Post reply on HN