Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

291–300 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#291
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

In Spotify’s defense, they used the pirated data only to show a proof of concept to the copyright holders, and that use was sanctioned by the local rights holders organization STIM.

The copyright holders then approved their concept, and subsequently Spotify got the rights to offer their service to customers. Everybody won.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#292
post #89

Earlier quoted context omitted.

The legal system is built to favor large corps and capital owners. See Katharina Pistor books for instance.

I think it’s the other way around. Those large entities break all the same laws and rules as others and then get to the point where they can influence the creation of a regulatory moat around themselves to prevent competitors from taking the same path as them.

I guess I sort of understand where this idea comes from, and when I was young I was totally into it, but now being in the corporate world for a decade and having my own small business, I just don't really see it anymore.

Big corps tend to be extremely conscientious of the the law. The law may not be ideal, but they tend to be hyper aware of it and have lawyers to ensure it. Small companies on the other hand are the wild fucking west, and tend to be overflowing with "turn a blind eye to that".

What big corps love is regulation that is expensive for small shops to overcome. They can drop $500k on a product cert no problem, be legally in the clear (and graciously compliant!), while making it near impossible for small guys to compete.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#294
post #202
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

The most outrageous thing about the whole story is that smart people (like here and not only) knew this all since day one. They been uncovering this the whole time. And in their face, with all the fierce ignorance, broligarchs deny, evade and totally pretend this never happened. The most non open company of all even went to lengths to accuse others of stealing their IP - not theirs to begin with. Just think of it - w…

Speaking of GPT2, I remember that nobody gave a shit what it was trained on, because it sucked then.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#295
post #186
post #144

Earlier quoted context omitted.

Exactly. Everyone should have the right to have access to this.

Are you sure that everything should be in the public domain? Say you spend a year writing a book, shouldn't you be able to sell it?

I have open sourced all work I legally can for the past 20 years, and it has only given me more exposure and made it easier for people to trust me with significant budget to solve their hardest problems.

Also I happily buy lots of books from people like Cory Doctorow and nostarchpress -because- their books are public and I want to support authors that value the freedom of their readers.

Books that are DRM or copyright protected however, I buy used paper copies or pirate because why would I financially support people that do not respect my freedom?

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#296

Earlier quoted context omitted.

> Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. I don't understand why you wouldn't just buy copies of the books. Seems like such a relatively inexpensive way to strengthen your legal case.

Buying a copy of the book doesn’t grant you the right to copy it. That is what copyright is for .

It grants you the right to read & study it though.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#298

I don't understand why it's even a question that Meta trained their LLM on copyrighted material. They say so in their paper! Quoting from their LLaMMa paper [Touvron et al., 2023]: > We include two book corpora in our training dataset: the Gutenberg Project, [...], and the Books3 section of ThePile (Gao et al., 2020), a publicly available dataset for training large language models. Following that reference: > Books3…

Critically, by torrenting they also directly distributed the copywritten material itself. That is a standalone infringement separate from any argument about trained LLMs.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#299
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

Everyone on here is smart enough. Just do not participate and save your money. Do not pay for digital goods. If Netflix raises their prices, it doesn't matter because there is a torrent of all of their shows. If Spotify raises their prices, it doesn't matter because your favorite artist has their entire library in a torrent. If some game company ask you to pay real life prices for a digital costume, find the crack on…

if you want to support an artist go to the show and BUY MERCH at the table! almost all of their income comes from that. the importance of buying a T-shirt at the show cannot be overstated and sometimes you get to say hi to your idol, too

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#300

Earlier quoted context omitted.

They should sell for a price it would make pirating it pointless. Like what Spotify or Netflix did to audio and visual content. Then they can use the exposure to find other ways to make money.

Or if you don't agree with the price, do not buy it. You are NOT entitled to entertain yourself in any way you want. (unless it's funded by taxes etc, in which case... okay, it's open to discussion.) Look, let's be honest - what gives you or others the right to steal from others?

If buying isn't owning, then piracy isn't stealing.
Post reply on HN