Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

281–290 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#281

Earlier quoted context omitted.

Wilhoit’s law: > There must be in-groups whom the law protects but does not bind, alongside out-groups whom the law binds but does not protect.

Is that a prescriptive or descriptive law?

He left out part of the quote, which is misappropriated as well. Wikipedia:

> This quotation is often incorrectly attributed to Francis M. Wilhoit:

> Conservatism consists of exactly one proposition, to wit: There must be in-groups whom the law protects but does not bind, alongside out-groups whom the law binds but does not protect.

> However, it was actually a 2018 blog response by 59-year-old Ohio composer Frank Wilhoit, years after Francis Wilhoit's death.

https://en.wikipedia.org/wiki/Francis_M._Wilhoit

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#282
post #193

Earlier quoted context omitted.

i know of a company that poisoned an entire town! thats terrorism if done by an individual. the company still exists, just paid a settlement and carried on...

I agree with your point, but will split hairs on using the word "terrorism". I think that should be reserved for people that commit atrocities for some political aim. I'm fairly sure the company in question (I assume Union Carbide) did not poison the town to advance a political agenda.

Profit at all costs is a political agenda.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#283

Earlier quoted context omitted.

> Thomas Babington Macaulay The one who got Hindu Sanskrit books translated in a horrible manner and then claimed: "I have no knowledge of either Sanskrit or Arabic. But I have done what I could to form a correct estimate of their value. I have read translations of the most celebrated Arabic and Sanskrit works. I have conversed both here and at home with men distinguished by their proficiency in the Eastern tongues.…

This is the corollary of the fallacy of appeal to authority: the rejection of an argument on the grounds that the speaker was horribly wrong on an unrelated or very loosely related topic. If you reject Macaulay on copyright because he was an imperialist, you can use the exact same logic to reject the arguments of essentially every person who ever lived. Very few humans who ever wrote anything important will perfectly…

I do think that context is still important in general, but probably only if you're doing deep research into Macaulay (or the specific target in mind). Treating everything in a vacuum isn't great either. Plenty of philosophical works for example, you really have to read in the time period and in the context of the author's life.

I find an acceptable tradeoff for now is, if I want to do deep research for myself, opening myself up to this sort of mushy subjective stuff is actually really important for making deep, objectively correct observations. Especially if the goal is to steelman, not strawman, the opponent's argument.

Otherwise, this kind of worst-case analysis thinking is fine. It's a logically sound conclusion, it's just kind of unsatisfying because we can't make stronger claims.

How do we decide when to make this tradeoff and for what things? Uhh.... idk. For me though, there has been value in using both kinds of thinking before though.

On a public forum, worst-case analysis is probably fine because the discussion ain't that deep. Also probably 90% of comments are made within the intention of a "gotcha" and not actually for discussion.

Basically, I totally agree with this, it's just that I've seen one too many online forums devolve into thought-terminating cliches using "rationality" as the basis. Here, I think it's totally justified to take this line... I instinctively had the same reaction upon reading GP's post (but then you could argue it's tone policing... and ahh we're off to the good ol' internet debate race spiral)

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#284
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

Yes, these companies are based on massive IP and copyright theft. And they still want to lecture others about their "property rights".

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#285
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

[dead]

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#286
post #166

Earlier quoted context omitted.

Airbnb and Uber have showed us that laws matter only to the extent that the political will to enforce them exists. Throw enough lawyers and lobbying money at the problem and the laws can simply be re-written to be friendlier to your business model.

The reason there was no political will to punish Airbnb and Uber for violating the law was that initially they were subsidized with VC money and so were able to undercut traditional hotels and taxis on price. In the world of tradable goods, pricing below cost with the intent of putting competition out of business so you can raise prices later is known as "dumping" and is itself illegal.

The traditional taxi industry was rife with corruption, bad experiences, and poor service in many jurisdictions before uber/lyft. As terrible of a human being that I think Travis Kalanick is, it was only going to take lawbreaking to overcome such a tainted system.

Medallion systems often prevented any competition, sometimes to absurd effect. The number of licenced taxis often didn't keep pace with population growth, sometimes even staying flat. Many drivers didn't own their own medallians then had to rent from owners, often making little money. In my city (Toronto) cabs were often dirty, broken, refused short distance fares (illegal) and smelled of cigarette smoke that was obviously from the driver.

Examples (paywalls, but you get the idea):

https://www.nytimes.com/1992/07/26/nyregion/amid-a-heritage-...

https://www.theglobeandmail.com/globe-drive/adventure/red-li...

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#287
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

Everyone on here is smart enough. Just do not participate and save your money. Do not pay for digital goods. If Netflix raises their prices, it doesn't matter because there is a torrent of all of their shows. If Spotify raises their prices, it doesn't matter because your favorite artist has their entire library in a torrent. If some game company ask you to pay real life prices for a digital costume, find the crack online and play on a private server. If YouTube wants to interrupt your video with an ad in the middle of the sentence, download one of the many options that blocks all ads. Billion dollar companies have shown they do not care about you. The people who complain about losing their salary, should just get replies thanking them for paying.

All the sad poor people who might be hurt were already paid. The caterer on your favorite show is not getting residuals. NBC also isn't going to stop making TV shows because that is all they can do. Content creators also existed on the internet long before that was a job. They just did it because they cared about it not for ad money. If you really want to support the artist directly go to a concert or just mail them a check. If you can't actually identify a person who might be hurt, then do not care.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#288

Earlier quoted context omitted.

I think it’s the other way around. Those large entities break all the same laws and rules as others and then get to the point where they can influence the creation of a regulatory moat around themselves to prevent competitors from taking the same path as them.

True but lets take examples one by one to see what we can learn : Spotify was doing illegal things until they made a deal to become legal and not to be trialed over what they done. Seems like business deals is what saved them, not regulatory capture (the regulations around IP for music pre existed Spotify)

Same would be the case for YouTube. Google case was different in that AFAIK there wasn't any obvious legal problem with indexing, and, back then, they were actually doing everyone a favor.

Hardly anyone had any issue with Google search until the time when news media screwed themselves over by going all in on ads, overdoing it, then trying to bring back the paywall, only to realize no one is actually browsing their sites but instead relies on Google to find specific articles. All kinds of legal and technical nonsense started happening (and then Google improved the blurbs under search results and added the "answer box", leading publishers big and small to collectively lose their minds...).

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#289
I don't understand why it's even a question that Meta trained their LLM on copyrighted material. They say so in their paper! Quoting from their LLaMMa paper [Touvron et al., 2023]:

> We include two book corpora in our training dataset: the Gutenberg Project, [...], and the Books3 section of ThePile (Gao et al., 2020), a publicly available dataset for training large language models.

Following that reference:

> Books3 is a dataset of books derived from a copy of the contents of the Bibliotik private tracker made available by Shawn Presser (Presser, 2020).

(Presser, 2020) refers to https://twitter.com/theshawwn/status/1320282149329784833. (Which funnily refers to this DMCA policy: https://the-eye.eu/dmca.mp4)

Furthermore, they state they trained on GitHub, web pages, and ArXiv, which are all contain copyrighted content.

Surely the question is: is it legal to train and/or use and/or distribute an AI model (or its weights, or its outputs) that is trained using copyrighted material. That it was trained on copyrighted material is certain.

[Touvron et al., 2023] https://arxiv.org/pdf/2302.13971

[Gao et al., 2020] https://arxiv.org/pdf/2101.00027

Post reply on HN