Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

361–370 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#361
post #305

> By September 2023, Bashlykov had seemingly dropped the emojis, consulting the legal team directly and emphasizing in an email that "using torrents would entail ‘seeding’ the files—i.e., sharing the content outside, this could be legally not OK." I'm pretty sure you can theoretically download torrents without seeding, although this is frowned upon. If they really seeded (with full bandwidth?) that's indeed pretty br…

You can, in transmission for example you can just set the seed percentage to 0%. I recognise that this makes me a bad torrenter, but I've been told in the past that my ISP wont be too happy about me seeding, and they already do something screwy to torrents I access through the surface web, so I'm just playing it safe

I think your client may still be sharing IP addresses, not sure about the legality of that

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#362

I don't understand why it's even a question that Meta trained their LLM on copyrighted material. They say so in their paper! Quoting from their LLaMMa paper [Touvron et al., 2023]: > We include two book corpora in our training dataset: the Gutenberg Project, [...], and the Books3 section of ThePile (Gao et al., 2020), a publicly available dataset for training large language models. Following that reference: > Books3…

Critically, by torrenting they also directly distributed the copywritten material itself. That is a standalone infringement separate from any argument about trained LLMs.

And punishing them in the normal manner will be an incredibly small slap on the wrist, and do absolutely nothing to help us find out what will play out in court regarding a fair-use defense on training AI with copyrighted material.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#363
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

[flagged]

For downloading papers, paid for with government funding and gatekept by greedy rent seekers charging ~ $30 a pop. The lengths people will go to defend things that should not exist astounds.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#364
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

> If you plug a laptop into a closet at MIT to download some scientific papers you forfeit your life.

In case anybody here doesn't know, that's a reference to Aaron Swartz, an activist (and Reddit co-founder) that was risking 35 years in prison and a $1 million fine just for downloading a lot of academic papers from JSTOR. He eventually took his life because of the pressure. May his soul rest in peace.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#365

It really makes you think about those crazy internet folks from back in the day who thought copyright law was too strict and that restricting humanity to knowledge in such a way was holding us all back for the benefit of a tiny few.

I'm all for chopping up copyright law. But until we do so, companies like Meta need to be treated just like everyone else. That means lawsuits, prison sentences, and millions in fines. And that's just the piracy part, there's also the lying/fraud part. Interestingly, a Dutch LLM project was sent a cease and desist after the local copyright lobby caught wind of it being trained on a bunch of pirated eBooks. The case u…

>need to be treated just like everyone else.

So a copyright warning letter in the mail from their ISP? Maybe someone should tell them about VPNs...

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#366
post #263
post #241

Earlier quoted context omitted.

…why? Will people buy less books because we have intuitive algorithms trained on old books? Personally, I strongly believe that the aesthetic skills of humanity are one of our most advanced faculties — we are nowhere close to replacing them with fully-automated output, AGI or no.

old books? i can imagine the shit/hallucinated-like generative AI we would have if the training weight was restricted to public domain stuff... i think when chatGPT was around version 2 or 3, i had extracted almost 2 pages (without any alteration from the original) with questions that considered the author from this book here, https://www.amazon.com/Loneliness-Human-Nature-Social-Connec... now it's up to you to think…

I find this such a strange remark on this front.

You got less than 1% of a book... from an author who has passed away... who wrote on a research topic that was funded by an institution that takes in hundreds of millions of dollars in federal grants each year...

I'm not an author (although I do generate almost exclusively IP for a living) and I think this is about as weak a form of this argument as you possibly make.

So right back at ya... who was hurt in your example?

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#367
post #58

We all like hating big corporations, especially Meta, and people seem to use this as an opportunity to advocate for punishing them. I think it's wiser to advocate for changing our IP laws.

We're sick of the double standards. https://en.wikipedia.org/wiki/Aaron_Swartz#United_States_v._... https://en.wikipedia.org/wiki/Aaron_Swartz#Death While Aaron Swartz was bullied to suicide, these corporations will walk free and make billions. I say give every tech CEO the Swartz treatment, then change the law.

The lesson here is make sure you only break the rules in the limits of severity that your wealth class allows.

MIT students will get away with breaking bigger rules than community college students will.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#368
post #294
post #202

Earlier quoted context omitted.

The most outrageous thing about the whole story is that smart people (like here and not only) knew this all since day one. They been uncovering this the whole time. And in their face, with all the fierce ignorance, broligarchs deny, evade and totally pretend this never happened. The most non open company of all even went to lengths to accuse others of stealing their IP - not theirs to begin with. Just think of it - w…

Speaking of GPT2, I remember that nobody gave a shit what it was trained on, because it sucked then.

[deleted]

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#369
post #325

Earlier quoted context omitted.

I don’t think this is true. At least in music, bands make far more money from touring and merch than they do from music sales. If copyright disappeared altogether, most smaller artists would be just fine because they have loyal fans and adjacent monetization strategies. See: Grateful Dead. They did just fine despite encouraging infringement of IP. IMO copyright mostly serves to protect the very biggest artists and co…

I think the point was that the big corporations get the money from selling music. And saying that bands currently make more money from touring kind of proves the point. They get too low % cut of music sales.

But the point of the response is that "getting money from selling music" is, in digital era, artificial scarcity. I.e. the copyright laws that big corporations are lobbying for continued enforcement and tightening, are the very thing that create this artificial scarcity that they are best positioned to profit off.

Cut out copyright, and no one will be getting any money from selling music per copy (or equivalent) - as it should be.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#370
post #342

Earlier quoted context omitted.

Hotels were just fine. Taxis were discriminatory and "uncool" to the point were Uber has saved thousands by preventing drunk driving. Now if you go out with the boys and get drunk, it's a 30 second casual call to get an Uber and get home. Live in a neighborhood Taxis are afraid to service,you can either make some extra income working for Uber or use it yourself. When Ubers used as its intended purpose, to basically m…

> Say your rents it's going to be late, you can pick up 20 or 30 hours of Uber this month to make it happen that sounds so incredibly dystopian, not sure if that was the intention :(

I forgot to mention, Jitney cabs( unlicensed cabs primarily serving minority neighborhoods)came long before Ubers. They're more of less gone now though.

What's better. Taking out a 300% APR payday loan, getting evicted or working an extra 20, 30 hours of Uber.

Post reply on HN