> In response, NVIDIA defended its actions as fair use, noting that books are nothing more than statistical correlations to its AI models. Does this even make sense? Are the copyright laws so bad that a statement like this would actually be in NVIDIA’s favor?
It does make sense. It’s controversial. Your memory memorizes things in the same way. So what nvidia does here is no different, the AI doesn’t actually copy any of the books. To call training illegal is similar to calling reading a book and remembering it illegal. Our copyright laws are nowhere near detailed enough to specify anything in detail here so there is indeed a logical and technical inconsistency here. I can…
Nvidia contacted Anna's Archive to access books
71–80 of 160 posts
Re: Nvidia contacted Anna's Archive to access books
#72Just to clarify, the most valuable company in the world refuses to pay for digital media?
Re: Nvidia contacted Anna's Archive to access books
#73Just to clarify, the most valuable company in the world refuses to pay for digital media?
I assume you're expecting that they'll reach out and cut a deal with each publishing house separately, and then those publishing houses will have to somehow transfer their data over to NVIDIA. But that's a very custom set of discussions and deals that have to be struck.
I think they're going to the pirate libraries because the product they want doesn't exist.
Re: Nvidia contacted Anna's Archive to access books
#74I'm wondering what Amazon is planning to do with their access to all those Kindle books.
• Anna’s Archive: ~61.7 million “books” (plus ~95.7M papers) as of January 2026 https://en.wikipedia.org/wiki/Anna%27s_Archive • Amazon Kindle: “over 6 million titles” as of March 2018 https://en.wikipedia.org/wiki/Anna%27s_Archive
Hard to compare because AA contains duplicates, and the Kindle number is old, but at a glance it seems AA wins.
Re: Nvidia contacted Anna's Archive to access books
#75Earlier quoted context omitted.
Models don’t reproduce books though. It’s impossible for a model to reproduce something word for word because the model never copied the book. Most of the best fit curve runs along a path that doesn’t even touch an actual data point.
If there is one exact sentence taken out of the book and not referenced in quotes and exact source, that triggers copyright laws. So model doesnt have to reproduce the entire book, it only required to reproduce one specific sentence (which may be a characteristic sentence to that author or to that book).
Yes, and that's stupid, and will need to be changed.
Re: Nvidia contacted Anna's Archive to access books
#76Earlier quoted context omitted.
In that case I would say it is the act of reproducing the books that is illegal. Training the AI on said books is not. So the illegality rests at the point of output and not at the point of input. I’m just speaking in terms of the technical interpretation of what’s in place. My personal views on what it should be are another topic.
> So the illegality rests at the point of output and not at the point of input. It's not as simple as that, as this settlement shows [1]. Also, generating output is what these models are primarily trained for. [1]: https://www.bbc.com/news/articles/c5y4jpg922qo
Yes but not generating illegal output. These models were trained with intent to generate legal output. The fact that it can generate illegal output is a side effect. That's my point.
If you use AI to generate illegal output, that act is illegal. If you use AI to generate legal output that act is not illegal. Thus the point of output is where the legal question lies. From inception up to training there is clear legal precedence for the existence of AI models.
Re: Nvidia contacted Anna's Archive to access books
#77Earlier quoted context omitted.
You can only read the book, if you purchased it. Even if you dont have the intent to reproduce it, you must purchase it. So, I guess NVDA should just purchase all those books, no?
Obviously not; one can borrow books from libraries and read them as well.
Re: Nvidia contacted Anna's Archive to access books
#78Earlier quoted context omitted.
It does make sense. It’s controversial. Your memory memorizes things in the same way. So what nvidia does here is no different, the AI doesn’t actually copy any of the books. To call training illegal is similar to calling reading a book and remembering it illegal. Our copyright laws are nowhere near detailed enough to specify anything in detail here so there is indeed a logical and technical inconsistency here. I can…
You need to pay for the books before you memorize them
The government is in full support of this "lending" concept, in fact they have created entire facilities devoted to this very concept of lending out books.
Re: Nvidia contacted Anna's Archive to access books
#79Just to clarify, the most valuable company in the world refuses to pay for digital media?
I see this sentiment posted quite a bit, but have the publishers made any products available that would allow AI training on their works for payment? A naive approach would be to go to an online bookstore and pay $15 for every book, but then you have copyrighted content that is encrypted, that it's a violation of the DMCA to decrypt. I assume you're expecting that they'll reach out and cut a deal with each publishing…
Re: Nvidia contacted Anna's Archive to access books
#80Earlier quoted context omitted.
> Does this even make sense? Are the copyright laws so bad that a statement like this would actually be in NVIDIA’s favor? It makes some sense, yeah. There's also precedent, in google scanning massive amounts of books, but not reproducing them. Most of our current copyright laws deal with reproductions. That's a no-no. It gets murky on the rest. Nvda's argument here is that they're not reproducing the works, they're…
Scanning books is literally reproducing them. Copying books from Anna's Archive is also literally reproducing them. The idea that it is only copyright infringement if you engage in further reproduction is just wrong. As a consumer you are unlikely to be targeted for such "end-user" infringement, but that doesn't mean it's not infringement.
This is the conclusion of the saga between the author's guild v. google. It goes through a lot of factors, but in the end the conclusion is this:
> In sum, we conclude that: (1) Google’s unauthorized digitizing of copyright-protected works, creation of a search functionality, and display of snippets from those works are non-infringing fair uses. The purpose of the copying is highly transformative, the public display of text is limited, and the revelations do not provide a significant market substitute for the protected aspects of the originals. Google’s commercial nature and profit motivation do not justify denial of fair use. (2) Google’s provision of digitized copies to the libraries that supplied the books, on the understanding that the libraries will use the copies in a manner consistent with the copyright law, also does not constitute infringement. Nor, on this record, is Google a contributory infringer.