Live data from Hacker News

Nvidia contacted Anna's Archive to access books

torrentfreak.com

111–120 of 160 posts

Re: Nvidia contacted Anna's Archive to access books

#111
post #3

> In response, NVIDIA defended its actions as fair use, noting that books are nothing more than statistical correlations to its AI models. Does this even make sense? Are the copyright laws so bad that a statement like this would actually be in NVIDIA’s favor?

Yes, it's been discussed many times before. All the corporations training LLMs have to have done a legal analysis and concluded that it's defensible. Even one of the white papers commissioned by the FSF ( " Copyright Implications of the Use of Code Repositories to Train a Machine Learning Model " at https://www.fsf.org/licensing/copilot/copyright-implications... ), concluded that using copyrighted data to train AI wa…

So it's legal to train an "intelligence" on everything for free based on fair use, but it's not legal to train another intelligence (my brain) on it?

Re: Nvidia contacted Anna's Archive to access books

#112
post #3

> In response, NVIDIA defended its actions as fair use, noting that books are nothing more than statistical correlations to its AI models. Does this even make sense? Are the copyright laws so bad that a statement like this would actually be in NVIDIA’s favor?

Copyright laws are so undefined and NVIDIAs lawyers so plentiful that the statement works in their favor. You're allowed to copy part of a work in many cases, the easiest example is you can quote a line from a book in a review. The line is fuzzy.

Re: Nvidia contacted Anna's Archive to access books

#113
post #68

Earlier quoted context omitted.

Indeed, the "copy" of the movie in your brain is not illegal. It would be rather troublesome and dystopian if it were.

The problem is when you use your "copy" as inspiration and actually create and publish something. It is very hard to be certain you are safe, besides literal expression close paraphrasing is also infringing, using world building elements, or using any original abstraction (AFC test). You can only know after a lawsuit. It is impossible to tell how much AI any creator used secretly, so now all works are under suspicion…

> close paraphrasing is also infringing, using world building elements, or using any original abstraction (AFC test)

World building elements? Do you have more details on that, because that feels wrong to me.

Unless you mean the specific names of things in the world like "Hobbits".

Re: Nvidia contacted Anna's Archive to access books

#114
post #90

Earlier quoted context omitted.

Did you pirated this movie? No I did not, it is fair use because this movie is nothing more than a statistical correlation to my dopamine production.

>Did you pirated this movie? No I did not, [...] You're probably being sarcastic but that's actually how the law works. You'll note that when people get sued for "pirating" movies, it's almost always because they were caught seeding a torrent, not for the act of watching an illegal copy. Movie studios don't go after visitors of illegal streaming sites, for instance.

> Movie studios don't go after visitors of illegal streaming sites, for instance.

They absolutely do, in France we have Hadopi that tracks torrent leecher. Hadopi had been heavily pushed by the movie and music industry.

Re: Nvidia contacted Anna's Archive to access books

#115

Earlier quoted context omitted.

> To call training illegal is similar to calling reading a book and remembering it illegal. Perhaps, but reproducing the book from this memory could very well be illegal. And these models are all about production.

Models don’t reproduce books though. It’s impossible for a model to reproduce something word for word because the model never copied the book. Most of the best fit curve runs along a path that doesn’t even touch an actual data point.

Models absolutely do reproduce books.

> With a simple two-phase procedure, we show that it is possible to extract large amounts of in-copyright text from four production LLMs. While we needed to jailbreak Claude 3.7 Sonnet and GPT-4.1 to facilitate extraction, Gemini 2.5 Pro and Grok 3 directly complied with text continuation requests. For Claude 3.7 Sonnet, we were able to extract four whole books near-verbatim, including two books under copyright in the U.S.: Harry Potter and the Sorcerer’s Stone and 1984.

https://arxiv.org/abs/2601.02671

Re: Nvidia contacted Anna's Archive to access books

#116
post #3

> In response, NVIDIA defended its actions as fair use, noting that books are nothing more than statistical correlations to its AI models. Does this even make sense? Are the copyright laws so bad that a statement like this would actually be in NVIDIA’s favor?

> Does this even make sense? Are the copyright laws so bad that a statement like this would actually be in NVIDIA’s favor? It makes some sense, yeah. There's also precedent, in google scanning massive amounts of books, but not reproducing them. Most of our current copyright laws deal with reproductions. That's a no-no. It gets murky on the rest. Nvda's argument here is that they're not reproducing the works, they're…

[deleted]

Re: Nvidia contacted Anna's Archive to access books

#117

Earlier quoted context omitted.

I see this sentiment posted quite a bit, but have the publishers made any products available that would allow AI training on their works for payment? A naive approach would be to go to an online bookstore and pay $15 for every book, but then you have copyrighted content that is encrypted, that it's a violation of the DMCA to decrypt. I assume you're expecting that they'll reach out and cut a deal with each publishing…

Do you believe in private property rights? If the product they want doesn't exist then they're shit out of luck and they must either make one or wait for one to get made. You're arguing that it's okay for them to break the law because doing business legally is really inconvenient . That would be the end of discussion if we lived in a world governed by the rule of law but we're repeatedly reminded that we don't.

Not arguing it's ok to break the law, but rather examining their incentives and alternatives, along with their associated costs.

Re: Nvidia contacted Anna's Archive to access books

#119

Earlier quoted context omitted.

I see this sentiment posted quite a bit, but have the publishers made any products available that would allow AI training on their works for payment? A naive approach would be to go to an online bookstore and pay $15 for every book, but then you have copyrighted content that is encrypted, that it's a violation of the DMCA to decrypt. I assume you're expecting that they'll reach out and cut a deal with each publishing…

> I assume you're expecting that they'll reach out and cut a deal with each publishing house separately, and then those publishing houses will have to somehow transfer their data over to NVIDIA. But that's a very custom set of discussions and deals that have to be struck. If this is the only legal way for them to train, then yes that is what they should do instead of breaking the law... just because its not easy does…

My comment is being misread as my support for piracy; my comment isn't meant to discuss anything at all about piracy. It's instead intended to look at everything that's not piracy, and examining their costs, and why the industry chose the path they did.

Existing rulings are beginning to suggest that if the books can be obtained legally, a separate license is not required for training. So I'm naturally interested in legal ways folks training models would get a lot of books, and whether the publishing industry has even considered the value there.

Re: Nvidia contacted Anna's Archive to access books

#120
post #106

Earlier quoted context omitted.

I see this sentiment posted quite a bit, but have the publishers made any products available that would allow AI training on their works for payment? A naive approach would be to go to an online bookstore and pay $15 for every book, but then you have copyrighted content that is encrypted, that it's a violation of the DMCA to decrypt. I assume you're expecting that they'll reach out and cut a deal with each publishing…

Hmm, didn't Anthropic buy a bunch of used books (like, physical ones), scanned them, and then destroyed them? If Anthropic can do that, surely can NVIDIA

Yes! And it was ruled legal by the courts, but the media spun it as "Anthropic destroys a million books to build AI". This is the only legal bulk approach I know of, hence my inquiry about such a product. I didn't expect such a harsh response from some of these comments.
Post reply on HN