Live data from Hacker News

Meta pirated at least 101 of my books, and others

garymarcus.substack.com

1–10 of 64 posts

Re: Meta pirated at least 101 of my books, and others

#3
(Shrug) We'll see what the courts say, Gary.

If training AI doesn't constitute fair use, you will lose more than you could ever possibly hope to gain. As will the rest of us.

Meanwhile, sublimate your dudgeon towards advocating for free access to the resulting models. That's what's important. Meta is not the company you want to go after here, since they released the resulting model weights.

Re: Meta pirated at least 101 of my books, and others

#5

(Shrug) We'll see what the courts say, Gary. If training AI doesn't constitute fair use, you will lose more than you could ever possibly hope to gain. As will the rest of us. Meanwhile, sublimate your dudgeon towards advocating for free access to the resulting models. That's what's important. Meta is not the company you want to go after here, since they released the resulting model weights.

Does fair use imply that pirating copyrighted material is ok?

I mean, it’s a serious question; I don’t see this as really connected.

As long as an AI can “understand” the content of a book and spit out a summary of it, or even leverage what it learned to perform further inference, I’d be inclined to say that this is fair use; a human would do the same.

But this has nothing to do with using pirated material for training, especially for some kind of commercial purpose (even if llama is free, they’re building on top of it) - I don’t see why it should be legal.

Re: Meta pirated at least 101 of my books, and others

#6
what would be interesting to learn is,

"Mega can regurgitate virtually any excerpt from any of my books, there for they have stolen them"

versus what is not interesting such as

"my books are in libgen therefore they stole my work, even though I can't find direct evidence of the theft"

>The most damning thing? It appears that Meta knew exactly what they are doing, and chose to proceed anyway.

that is not the most damning thing. It might trigger worse damages or elevation of the severity of an infraction, but it is not evidence of guilt per se, which is what I would call "damning"

Re: Meta pirated at least 101 of my books, and others

#8

(Shrug) We'll see what the courts say, Gary. If training AI doesn't constitute fair use, you will lose more than you could ever possibly hope to gain. As will the rest of us. Meanwhile, sublimate your dudgeon towards advocating for free access to the resulting models. That's what's important. Meta is not the company you want to go after here, since they released the resulting model weights.

Why should it be fair use? Why would being a derivative work not be OK? There is a massive corpus of public domain and FOSS works. Likewise plenty of permissively licensed government created datasets. There is no reason why any corpus created from these sources is insufficient.

Re: Meta pirated at least 101 of my books, and others

#9

(Shrug) We'll see what the courts say, Gary. If training AI doesn't constitute fair use, you will lose more than you could ever possibly hope to gain. As will the rest of us. Meanwhile, sublimate your dudgeon towards advocating for free access to the resulting models. That's what's important. Meta is not the company you want to go after here, since they released the resulting model weights.

Does fair use imply that pirating copyrighted material is ok? I mean, it’s a serious question; I don’t see this as really connected. As long as an AI can “understand” the content of a book and spit out a summary of it, or even leverage what it learned to perform further inference, I’d be inclined to say that this is fair use; a human would do the same. But this has nothing to do with using pirated material for traini…

I get the commercial/legal angle, but from the viewpoint of AI being something we as a society have an interest in developing, how should this work?

Do you want to severely limit evolution of models by having them pick (and buy) a tiny subset of all books?

Should every training run put money into a pool that gets paid out to every rights holder of every book that has ever been published?

Should Meta buy a physical or electronic copy of every book they want to use for training? That has zero impact on revenue for individual authors.

Would they be paid by word, by token, by book? This makes little sense. We don’t charge people for the knowledge they acquired while going to the library over 50 years, AI just squeezes this into weeks. Our legal framework simply doesn’t fit.

Re: Meta pirated at least 101 of my books, and others

#10

(Shrug) We'll see what the courts say, Gary. If training AI doesn't constitute fair use, you will lose more than you could ever possibly hope to gain. As will the rest of us. Meanwhile, sublimate your dudgeon towards advocating for free access to the resulting models. That's what's important. Meta is not the company you want to go after here, since they released the resulting model weights.

To point out the obvious.

Unauthorized copying (aka pirating) is definitely a copyright violation.

That appears to be a huge problem with the large models and training. They don't secure legal access to the materials they train on, and thus fail to compensate authors for their work.

AKA students are required to buy or otherwise obtain legal access to their text books(like checking the book out of the library).

Training AI should play the same rules humans students have to follow.

Post reply on HN