Meta claims torrenting pirated books isn't illegal without proof of seeding
71–80 of 471 posts
Re: Meta claims torrenting pirated books isn't illegal without proof of seeding
#72Earlier quoted context omitted.
I don't think that what Aaron did was* wrong. Meta's wholescale theft, however, is pretty hard to defend, and Meta knew it. That's why they went to some lengths to hide it. Similarly, that OpenAI whistleblower, the one whose family was calling for a murder investigation, might be alive today if it wasn't pretty well known that stealing the work of thousands/millions of people to make a for-profit imitation machine is…
What Aaron did was not wrong. He intended to make journal articles publicly available. They should be, as many are publicly funded, and academic publishers like Elsevier do not pay for these articles. Scientists provide them to journals. Universities, libraries, and we then have to buy back access.
Re: Meta claims torrenting pirated books isn't illegal without proof of seeding
#73Earlier quoted context omitted.
They can just bribe the president.
Maybe Trump will legalise internet-based copying. After all, the main people hurt would be Hollywood, which is run by people supporting the Democrats. And it would be popular with many voters (not an issue for Trump but it is for Republicans).
Counter example: ownership of Amazon MGM Studios and its parent Amazon.
Re: Meta claims torrenting pirated books isn't illegal without proof of seeding
#74Earlier quoted context omitted.
Paying the price of one copy does not imply that you can use it for training, right ?
(note; not a lawyer) It depends on if a model is a derivate work from it's source material or not. If yes, then all copyright protections come into force. If not, then the author can't rely on copyright to protect themselves. My instinct/gut says that an AI model is a derivative work from the training data (in that it quite literally takes training data to produce a new creative output, with the "human addition" bein…
Re: Meta claims torrenting pirated books isn't illegal without proof of seeding
#75It's certainly nice to see someone accused of bittorrenting with the bankroll to come up with a decent legal defense team.
The unfortunate side effect is that a megacorp gets to vacuum up the sum of human knowledge for free, boil it down, and sell it back to us for a nice profit.
Re: Meta claims torrenting pirated books isn't illegal without proof of seeding
#76Re: Meta claims torrenting pirated books isn't illegal without proof of seeding
#77"If you steal from one author, it's plagiarism; if you steal from many, it's research." - Wilson Mizner
"Plagiarize, let no one else's work evade your eyes. Remember why the good lord made your eyes, so don't shade your eyes but plagiarize, plagiarize, plagiarize !" ~ Me
Re: Meta claims torrenting pirated books isn't illegal without proof of seeding
#78Are they actually claiming only that they didn't share after the torrent completed? Or is the journalist just confused?
My understanding with bittorrent is that normally during download you are also uploading. "Seeding" is just what the uploading part is called when you're not also downloading.
I think it is possible to download without doing any uploading at all, but I feel like the onus of proof should be on them to show that they actually did that.
Re: Meta claims torrenting pirated books isn't illegal without proof of seeding
#79Earlier quoted context omitted.
Paying the price of one copy does not imply that you can use it for training, right ?
(note; not a lawyer) It depends on if a model is a derivate work from it's source material or not. If yes, then all copyright protections come into force. If not, then the author can't rely on copyright to protect themselves. My instinct/gut says that an AI model is a derivative work from the training data (in that it quite literally takes training data to produce a new creative output, with the "human addition" bein…
But it's not just supervised training. Maybe a model trained on reasoning traces and RLHF is not a mere derivative of the training set. All recent models are being trained on self generated data produced with reward or preference models.
When a model trains on a piece of text it won't derive gradients from the parts it knows, it will only absorb the novel parts. So what it takes from each example depends on the ordering of training examples. It is a process of diffing between model and text, could be seen as a form of analysis not simple memorization.
Even if it is infringement to train on protected works, the model size is 100x up to 1000x smaller than the training set, it has no space to memorize it.
The larger the training set, the less impact any one work has. It is de minimis use, paradoxically, the more you take the less you imitate.
That should matter when estimating damages.
Re: Meta claims torrenting pirated books isn't illegal without proof of seeding
#80Earlier quoted context omitted.
"Plagiarize, let no one else's work evade your eyes. Remember why the good lord made your eyes, so don't shade your eyes but plagiarize, plagiarize, plagiarize !" ~ Me
Glad to see you didn't acknowledge your source! :-) (It's Tom Lehrer, for any who don't recognize it.)