Live data from Hacker News

Meta pirated at least 101 of my books, and others

garymarcus.substack.com

21–30 of 64 posts

Re: Meta pirated at least 101 of my books, and others

#21

(Shrug) We'll see what the courts say, Gary. If training AI doesn't constitute fair use, you will lose more than you could ever possibly hope to gain. As will the rest of us. Meanwhile, sublimate your dudgeon towards advocating for free access to the resulting models. That's what's important. Meta is not the company you want to go after here, since they released the resulting model weights.

Why should it be fair use? Why would being a derivative work not be OK? There is a massive corpus of public domain and FOSS works. Likewise plenty of permissively licensed government created datasets. There is no reason why any corpus created from these sources is insufficient.

> Why would being a derivative work not be OK?

That's not even the real problem. It's a problem, yes, but not the real problem. The problem is that before they could train the model on the book, they had to copy the book from somewhere. Is it ok to make illegal pirated copies of a copyrighted book to train your model? I think that's the issue we are dealing with here.

Whether it is ok to create a derivative work or not is beside the point.

Re: Meta pirated at least 101 of my books, and others

#22
post #11

The idea that you can’t train on copyrighted materials is ludicrous, imho. So apparently you don’t want humanity and the future of intelligence to benefit from your work? You just want it to keep it locked up in some archive that virtually no one ever reads? Might as well say the people who read your books aren’t allowed to teach the concepts or theories. Completely asinine argument. If you don’t want the knowledge t…

[flagged]

Re: Meta pirated at least 101 of my books, and others

#23
post #11

The idea that you can’t train on copyrighted materials is ludicrous, imho. So apparently you don’t want humanity and the future of intelligence to benefit from your work? You just want it to keep it locked up in some archive that virtually no one ever reads? Might as well say the people who read your books aren’t allowed to teach the concepts or theories. Completely asinine argument. If you don’t want the knowledge t…

> The idea that you can’t train on copyrighted materials is ludicrous, imho.

Let us for a minute accept that it is ok to train on copyrighted materials. I don't believe that but I'll humor you. So let's accept it.

To train on copyrighted materials, they need to purchase the copyrighted materials, correct? If you wanted to train a model on all O'reilly books, you'd purchase the O'reilly books first, wouldn't you?

Do you think it is ok to make illegal pirated copies of the book to do your training?

Re: Meta pirated at least 101 of my books, and others

#24
Speaking broadly, the publishers who hold the copyrights on these materials have often behaved poorly. From overbroad DMCA takedown demands that chill fair use, to threatening libraries, students, and scholars with lawsuits and stiff penalties for minor infringements, to "copyright trolling" campaigns sending mass settlement demands to alleged infringers -- I have little sympathy for copyright holders.

I'm still angry about how publishers and the Authors Guild sued Google over Google Books. Intellectual property is why we as a society can't have nice things. While I'm not a fan of Meta, their open weight models are probably one of the best things they've ever done, and I'll back big tech over publishers every time.

Re: Meta pirated at least 101 of my books, and others

#25
post #11

The idea that you can’t train on copyrighted materials is ludicrous, imho. So apparently you don’t want humanity and the future of intelligence to benefit from your work? You just want it to keep it locked up in some archive that virtually no one ever reads? Might as well say the people who read your books aren’t allowed to teach the concepts or theories. Completely asinine argument. If you don’t want the knowledge t…

I’m not sure that’s the whole picture. Followed to the logical conclusion, everyone should have the right to pirate whatever books they want and then feed them into a local LLM. Which leads to less kickback to the author, which means they can’t sustainably write, and we end up in a worse off situation.

Re: Meta pirated at least 101 of my books, and others

#26
i really dont get this and i personally beleive the world would be a better place without IP in any form.

But also no one is selling "your book", the product is completely different in literally every conceivable way.

you have never (and no one ever should) own words arranged in a certain way. You own the right to sell a book. Not the words themselves.

meta does bad things and im not a fan, but this really pales in comparison.

Re: Meta pirated at least 101 of my books, and others

#27

i really dont get this and i personally beleive the world would be a better place without IP in any form. But also no one is selling "your book", the product is completely different in literally every conceivable way. you have never (and no one ever should) own words arranged in a certain way. You own the right to sell a book. Not the words themselves. meta does bad things and im not a fan, but this really pales in c…

Well, if you own copyright for a song, you can claim licensing fees for any public performance of those lyrics.

I wonder if an equivalent to Performance Rights Organizations will emerge as a channel for LLM publishers (so to speak) to pay fees.

Re: Meta pirated at least 101 of my books, and others

#28

Earlier quoted context omitted.

Does fair use imply that pirating copyrighted material is ok? I mean, it’s a serious question; I don’t see this as really connected. As long as an AI can “understand” the content of a book and spit out a summary of it, or even leverage what it learned to perform further inference, I’d be inclined to say that this is fair use; a human would do the same. But this has nothing to do with using pirated material for traini…

I get the commercial/legal angle, but from the viewpoint of AI being something we as a society have an interest in developing, how should this work? Do you want to severely limit evolution of models by having them pick (and buy) a tiny subset of all books? Should every training run put money into a pool that gets paid out to every rights holder of every book that has ever been published ? Should Meta buy a physical o…

> Should every training run put money into a pool that gets paid out to every rights holder of every book that has ever been published?

That could actually work. Bearing in mind that all copyright laws are messy and terrible, this proposal is at least not impossible.

"Ever been published" means in the last 100 years.

Re: Meta pirated at least 101 of my books, and others

#29

i really dont get this and i personally beleive the world would be a better place without IP in any form. But also no one is selling "your book", the product is completely different in literally every conceivable way. you have never (and no one ever should) own words arranged in a certain way. You own the right to sell a book. Not the words themselves. meta does bad things and im not a fan, but this really pales in c…

Well, if you own copyright for a song, you can claim licensing fees for any public performance of those lyrics. I wonder if an equivalent to Performance Rights Organizations will emerge as a channel for LLM publishers (so to speak) to pay fees.

and that is equally atrocious and should be eliminated from a society that wants to share ideas freely.

Idk if your in the US but you also massively oversimplify in your example, copyright law is waaaaaaaay more complex than that and it would take a set of special circumstances way beyond doing what you say it siphon money from an infringment claim

Re: Meta pirated at least 101 of my books, and others

#30

Earlier quoted context omitted.

Well, if you own copyright for a song, you can claim licensing fees for any public performance of those lyrics. I wonder if an equivalent to Performance Rights Organizations will emerge as a channel for LLM publishers (so to speak) to pay fees.

and that is equally atrocious and should be eliminated from a society that wants to share ideas freely. Idk if your in the US but you also massively oversimplify in your example, copyright law is waaaaaaaay more complex than that and it would take a set of special circumstances way beyond doing what you say it siphon money from an infringment claim

Yes, of course. I struggle to have an opinion here since I don't like either side in this fight, but I eventually squeezed one out, and there it is.
Post reply on HN