Live data from Hacker News

Nvidia contacted Anna's Archive to access books

torrentfreak.com

91–100 of 160 posts

Re: Nvidia contacted Anna's Archive to access books

#91

Earlier quoted context omitted.

In that case I would say it is the act of reproducing the books that is illegal. Training the AI on said books is not. So the illegality rests at the point of output and not at the point of input. I’m just speaking in terms of the technical interpretation of what’s in place. My personal views on what it should be are another topic.

> So the illegality rests at the point of output and not at the point of input. It's not as simple as that, as this settlement shows [1]. Also, generating output is what these models are primarily trained for. [1]: https://www.bbc.com/news/articles/c5y4jpg922qo

Unfortunately a settlement doesn't really show you anything definitive about the legality or illegality of something.

It only shows you that the defendant thought it would be better for them to pay up rather than continue to be dragged through court, and that the plaintiff preferred some amount of certain money now over some other amount of uncertain money later, or never.

We cannot say with any amount of confidence how the court would have ruled on the legality, had things been allowed to play out without a settlement.

Re: Nvidia contacted Anna's Archive to access books

#92

Earlier quoted context omitted.

Scanning books is literally reproducing them. Copying books from Anna's Archive is also literally reproducing them. The idea that it is only copyright infringement if you engage in further reproduction is just wrong. As a consumer you are unlikely to be targeted for such "end-user" infringement, but that doesn't mean it's not infringement.

Private reproductions are allowed (e.g. backups). Distributing them non-privately is not.

>Distributing them non-privately is not.

You can even distribute them, to some limits.

https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,....

Re: Nvidia contacted Anna's Archive to access books

#93

Earlier quoted context omitted.

> To call training illegal is similar to calling reading a book and remembering it illegal. A type of wishful thinking fallacy. In law scale matters. It's legal for you to possess a single joint. It's not legal to possess 400 tons of weed in a warehouse.

It is not the scale that matters here, in your example, but intent. With 1 joint, you want to smoke yourself. With 400, you very possibly want to sell it to others. Scale in itself doesnt matter, scale matters only as to the extent it changes what your intention may be.

Right, but in the weed analogy, the scale is used as a proxy to assume intent. When someone is caught with those 400 joints, the prosecution doesn't have to prove intent, because the law has that baked in already.

You could say the same in LLM training, that doing so at scale implies the intent to commit copyright infringement, whereas reading a single book does not. (I don't believe our current law would see it this way, but it wouldn't be inconsistent if it did, or if new law would be written to make it so.)

Re: Nvidia contacted Anna's Archive to access books

#95

Earlier quoted context omitted.

It is not the scale that matters here, in your example, but intent. With 1 joint, you want to smoke yourself. With 400, you very possibly want to sell it to others. Scale in itself doesnt matter, scale matters only as to the extent it changes what your intention may be.

It’s clear nvidia and every single one of these big AI corps do not want their AIs to violate the law. The intent is clear as day here. Scale is only used for emergence, openAI found that training transformers on the entire internet would make is more then just a next token predictor and that is the intent everyone is going for when building these things.

I don't think that's clear at all. Businesses routinely break the law if they believe the benefits in doing so will outweigh the consequences.

I think this is even more common and more brazen when it comes to "disruptive" businesses and technologies.

Re: Nvidia contacted Anna's Archive to access books

#96

Earlier quoted context omitted.

> To call training illegal is similar to calling reading a book and remembering it illegal. A type of wishful thinking fallacy. In law scale matters. It's legal for you to possess a single joint. It's not legal to possess 400 tons of weed in a warehouse.

Er no. I’ve read and remember hundreds of books in my life time. It’s not any more illegal based off scale. The law doesn’t differentiate whether I remember one book or a hundred then there’s no difference for thousands or millions. No wishful thinking here.

What is "scale" in this context? I think arguably 100 books over the span of decades is not "scale".

But tens (hundreds?) of thousands of books over the span of a few weeks? That's definitely "scale".

Re: Nvidia contacted Anna's Archive to access books

#97

Earlier quoted context omitted.

Obviously not; one can borrow books from libraries and read them as well.

That's true. But the book itself was legally purchased. So if nvidia went to the library and trained AI by borrowing books, that should be technically legal.

Do you have the same legal rights to something that you've borrowed as you do with something you've purchased, though?

Would it be legal for me to borrow a book from the library, then scan and OCR every page and create an EPUB file of the result? Even if I didn't distribute it, that sounds questionable to me. Whereas if I had purchased the book and done the same, I believe that might be ok (format shifting for personal use).

Back when VHS and video rental was a thing, my parents would routinely copy rented VHS tapes if we liked the movie (camcorder connected to VCR with composite video and audio cables, worked great if there wasn't Macrovision copy protection on the source). I don't think they were under any illusions that what they were doing was ok.

Re: Nvidia contacted Anna's Archive to access books

#98

Just to clarify, the most valuable company in the world refuses to pay for digital media?

I see this sentiment posted quite a bit, but have the publishers made any products available that would allow AI training on their works for payment? A naive approach would be to go to an online bookstore and pay $15 for every book, but then you have copyrighted content that is encrypted, that it's a violation of the DMCA to decrypt. I assume you're expecting that they'll reach out and cut a deal with each publishing…

> I assume you're expecting that they'll reach out and cut a deal with each publishing house separately, and then those publishing houses will have to somehow transfer their data over to NVIDIA. But that's a very custom set of discussions and deals that have to be struck.

If this is the only legal way for them to train, then yes that is what they should do instead of breaking the law... just because its not easy doesn't mean piracy is fine.

Re: Nvidia contacted Anna's Archive to access books

#99

Just to clarify, the most valuable company in the world refuses to pay for digital media?

I see this sentiment posted quite a bit, but have the publishers made any products available that would allow AI training on their works for payment? A naive approach would be to go to an online bookstore and pay $15 for every book, but then you have copyrighted content that is encrypted, that it's a violation of the DMCA to decrypt. I assume you're expecting that they'll reach out and cut a deal with each publishing…

That's not relevant went it comes to copyright law. The copyright holder has the sole legal right to decide how the work is distributed.

If it isn't distributed in a manner to your liking, the only legal thing you can do is not have a copy of it at all.

Re: Nvidia contacted Anna's Archive to access books

#100
post #97

Earlier quoted context omitted.

That's true. But the book itself was legally purchased. So if nvidia went to the library and trained AI by borrowing books, that should be technically legal.

Do you have the same legal rights to something that you've borrowed as you do with something you've purchased, though? Would it be legal for me to borrow a book from the library, then scan and OCR every page and create an EPUB file of the result? Even if I didn't distribute it, that sounds questionable to me. Whereas if I had purchased the book and done the same, I believe that might be ok (format shifting for person…

Well If I copied it word for word maybe, but if I read it and "trained it" into my brain then it's clearly not illegal.

SO the grey area here is if I "trained" an LLM in a similar way and not copied it word for word then is it legal? Because fundamentally speaking it's literally the same action taken.

Post reply on HN