> By September 2023, Bashlykov had seemingly dropped the emojis, consulting the legal team directly and emphasizing in an email that "using torrents would entail ‘seeding’ the files—i.e., sharing the content outside, this could be legally not OK." I'm pretty sure you can theoretically download torrents without seeding, although this is frowned upon. If they really seeded (with full bandwidth?) that's indeed pretty br…
You can, in transmission for example you can just set the seed percentage to 0%. I recognise that this makes me a bad torrenter, but I've been told in the past that my ISP wont be too happy about me seeding, and they already do something screwy to torrents I access through the surface web, so I'm just playing it safe
Meta torrented & seeded 81.7 TB dataset containing copyrighted data
361–370 of 981 posts
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#362I don't understand why it's even a question that Meta trained their LLM on copyrighted material. They say so in their paper! Quoting from their LLaMMa paper [Touvron et al., 2023]: > We include two book corpora in our training dataset: the Gutenberg Project, [...], and the Books3 section of ThePile (Gao et al., 2020), a publicly available dataset for training large language models. Following that reference: > Books3…
Critically, by torrenting they also directly distributed the copywritten material itself. That is a standalone infringement separate from any argument about trained LLMs.
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#363Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…
[flagged]
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#364Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…
In case anybody here doesn't know, that's a reference to Aaron Swartz, an activist (and Reddit co-founder) that was risking 35 years in prison and a $1 million fine just for downloading a lot of academic papers from JSTOR. He eventually took his life because of the pressure. May his soul rest in peace.
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#365It really makes you think about those crazy internet folks from back in the day who thought copyright law was too strict and that restricting humanity to knowledge in such a way was holding us all back for the benefit of a tiny few.
I'm all for chopping up copyright law. But until we do so, companies like Meta need to be treated just like everyone else. That means lawsuits, prison sentences, and millions in fines. And that's just the piracy part, there's also the lying/fraud part. Interestingly, a Dutch LLM project was sent a cease and desist after the local copyright lobby caught wind of it being trained on a bunch of pirated eBooks. The case u…
So a copyright warning letter in the mail from their ISP? Maybe someone should tell them about VPNs...
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#366Earlier quoted context omitted.
…why? Will people buy less books because we have intuitive algorithms trained on old books? Personally, I strongly believe that the aesthetic skills of humanity are one of our most advanced faculties — we are nowhere close to replacing them with fully-automated output, AGI or no.
old books? i can imagine the shit/hallucinated-like generative AI we would have if the training weight was restricted to public domain stuff... i think when chatGPT was around version 2 or 3, i had extracted almost 2 pages (without any alteration from the original) with questions that considered the author from this book here, https://www.amazon.com/Loneliness-Human-Nature-Social-Connec... now it's up to you to think…
You got less than 1% of a book... from an author who has passed away... who wrote on a research topic that was funded by an institution that takes in hundreds of millions of dollars in federal grants each year...
I'm not an author (although I do generate almost exclusively IP for a living) and I think this is about as weak a form of this argument as you possibly make.
So right back at ya... who was hurt in your example?
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#367We all like hating big corporations, especially Meta, and people seem to use this as an opportunity to advocate for punishing them. I think it's wiser to advocate for changing our IP laws.
We're sick of the double standards. https://en.wikipedia.org/wiki/Aaron_Swartz#United_States_v._... https://en.wikipedia.org/wiki/Aaron_Swartz#Death While Aaron Swartz was bullied to suicide, these corporations will walk free and make billions. I say give every tech CEO the Swartz treatment, then change the law.
MIT students will get away with breaking bigger rules than community college students will.
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#368Earlier quoted context omitted.
The most outrageous thing about the whole story is that smart people (like here and not only) knew this all since day one. They been uncovering this the whole time. And in their face, with all the fierce ignorance, broligarchs deny, evade and totally pretend this never happened. The most non open company of all even went to lengths to accuse others of stealing their IP - not theirs to begin with. Just think of it - w…
Speaking of GPT2, I remember that nobody gave a shit what it was trained on, because it sucked then.
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#369Earlier quoted context omitted.
I don’t think this is true. At least in music, bands make far more money from touring and merch than they do from music sales. If copyright disappeared altogether, most smaller artists would be just fine because they have loyal fans and adjacent monetization strategies. See: Grateful Dead. They did just fine despite encouraging infringement of IP. IMO copyright mostly serves to protect the very biggest artists and co…
I think the point was that the big corporations get the money from selling music. And saying that bands currently make more money from touring kind of proves the point. They get too low % cut of music sales.
Cut out copyright, and no one will be getting any money from selling music per copy (or equivalent) - as it should be.
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#370Earlier quoted context omitted.
Hotels were just fine. Taxis were discriminatory and "uncool" to the point were Uber has saved thousands by preventing drunk driving. Now if you go out with the boys and get drunk, it's a 30 second casual call to get an Uber and get home. Live in a neighborhood Taxis are afraid to service,you can either make some extra income working for Uber or use it yourself. When Ubers used as its intended purpose, to basically m…
> Say your rents it's going to be late, you can pick up 20 or 30 hours of Uber this month to make it happen that sounds so incredibly dystopian, not sure if that was the intention :(
What's better. Taking out a 300% APR payday loan, getting evicted or working an extra 20, 30 hours of Uber.