Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

301–310 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#301
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

Everyone on here is smart enough. Just do not participate and save your money. Do not pay for digital goods. If Netflix raises their prices, it doesn't matter because there is a torrent of all of their shows. If Spotify raises their prices, it doesn't matter because your favorite artist has their entire library in a torrent. If some game company ask you to pay real life prices for a digital costume, find the crack on…

I'm not paying for Led Zeppelin IV after having probably bought 3 copies in my lifetime. I agree with you.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#302
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

" If you plug a laptop into a closet at MIT to download some scientific papers you forfeit your life." This is exactly what I immediately thought while reading the article. It almost feels like the legal system only punishes general public, while most of these guys are above it.

At this point, I think it's safe to say it doesn't 'feel' that way. It is that way. Sorry if you were being facetious and I didn't pick up on it.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#303
And they're going to get away with it simply because if you or I openly did this the DMCA fines would be for a million trillion dollars. Since Meta shareholders can't stomach a million trillion dollars in fines, their lawyers will wave their magic wands and poof! No laws were broken!

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#304
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

" If you plug a laptop into a closet at MIT to download some scientific papers you forfeit your life." This is exactly what I immediately thought while reading the article. It almost feels like the legal system only punishes general public, while most of these guys are above it.

It is more a money thing. Meta can pay x billion like pocket change. Regular people are run through the ringer to teach the plebs to not get out of line.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#305
> By September 2023, Bashlykov had seemingly dropped the emojis, consulting the legal team directly and emphasizing in an email that "using torrents would entail ‘seeding’ the files—i.e., sharing the content outside, this could be legally not OK."

I'm pretty sure you can theoretically download torrents without seeding, although this is frowned upon. If they really seeded (with full bandwidth?) that's indeed pretty brazen.

It is sort of strange that Meta is being singled out here though, and sort of sad considering they at least release the model weights. What's the signal? Do illegal shit to be competitive, but make sure there is no evidence?

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#306
post #166

Earlier quoted context omitted.

Airbnb and Uber have showed us that laws matter only to the extent that the political will to enforce them exists. Throw enough lawyers and lobbying money at the problem and the laws can simply be re-written to be friendlier to your business model.

The hotel and taxi industry were legit terrible before those two disrupted them. Laws are ment to be broken. Especially in cronist systems where incumbents write the laws.

Not terrible everywhere.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#307

Earlier quoted context omitted.

> Thomas Babington Macaulay The one who got Hindu Sanskrit books translated in a horrible manner and then claimed: "I have no knowledge of either Sanskrit or Arabic. But I have done what I could to form a correct estimate of their value. I have read translations of the most celebrated Arabic and Sanskrit works. I have conversed both here and at home with men distinguished by their proficiency in the Eastern tongues.…

This is the corollary of the fallacy of appeal to authority: the rejection of an argument on the grounds that the speaker was horribly wrong on an unrelated or very loosely related topic. If you reject Macaulay on copyright because he was an imperialist, you can use the exact same logic to reject the arguments of essentially every person who ever lived. Very few humans who ever wrote anything important will perfectly…

> If you reject Macaulay on copyright because he was an imperialist

On the contrary I would argue that this is precisely why you SHOULD NOT take his opinion on copyright. One of the main outcomes of imperialism/colonization is denigrating/destroying/appropriating works of art, literature with the primary goal of subjugation, subversion and thereupon replacement of native culture/traditions/institutions. I did not quote the other half of his nauseating take but I'll post it nevertheless:

"[...] And I certainly never met with any Orientalist who ventured to maintain that the Arabic and Sanscrit poetry could be compared to that of the great European nations. But when we pass from works of imagination to works in which facts are recorded, and general principles investigated, the superiority of the Europeans becomes absolutely immeasurable. It is, I believe, no exaggeration to say, that all the historical information which has been collected from all the books written in the Sanscrit language is less valuable than what may be found in the most paltry abridgments used at preparatory schools in England. In every branch of physical or moral philosophy, the relative position of the two nations is nearly the same."

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#308

Earlier quoted context omitted.

The reason there was no political will to punish Airbnb and Uber for violating the law was that initially they were subsidized with VC money and so were able to undercut traditional hotels and taxis on price. In the world of tradable goods, pricing below cost with the intent of putting competition out of business so you can raise prices later is known as "dumping" and is itself illegal.

The traditional taxi industry was rife with corruption, bad experiences, and poor service in many jurisdictions before uber/lyft. As terrible of a human being that I think Travis Kalanick is, it was only going to take lawbreaking to overcome such a tainted system. Medallion systems often prevented any competition, sometimes to absurd effect. The number of licenced taxis often didn't keep pace with population growth,…

I've never been a huge user of either, but my worst Uber ride was much better than my best taxi ride.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#309

I don't understand why it's even a question that Meta trained their LLM on copyrighted material. They say so in their paper! Quoting from their LLaMMa paper [Touvron et al., 2023]: > We include two book corpora in our training dataset: the Gutenberg Project, [...], and the Books3 section of ThePile (Gao et al., 2020), a publicly available dataset for training large language models. Following that reference: > Books3…

There are two different things when it comes to discussing training LLM's on "copyright" protected data, and I almost never see people differentiate.

1.) Training on copyright that is publicly available. You write a poem and publish it online for the world to read. That is your IP, no one else can take it an sell it, but they are free to read and be inspired by it. The legalitly of training on this is in the courts, but so far seems to be going in favor of LLMs.

2.) Training on copyright that is not publicly available. These are pretty much pirated works or works obtained by backdoor to avoid paying for them. Your poem is behind a paywall and you never got paid, yet the poem is known by the LLM. This is just straight illegal, as you legally must pay to view the work. However there might be conditions here too like paying for access to an archive and then training on everything in it.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#310
post #222

I strongly urge people to read Thomas Babington Macaulay's speeches on copyright, its aims, terms, and hazards. Very well reasoned and explained. In particular, people often cited the case of authors who had died leaving a family in destitution, and claimed that copyright extension would be a fair way of preventing this, but in most cases the remaining family had never held the copyright; the author had initally sold…

I’m a huge IP hater and am sure that happens, but to be fair, letting copyright extend past death also increases the amount the author can sell it for in the first place.

The current workaround is to attribute footnotes to your beneficiaries, or quote them in the dedication. Those become derivative works subject to the lifetime of your beneficiary.
Post reply on HN