Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

421–430 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#421
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

[flagged]

What a repulsive comment.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#422
post #339

Earlier quoted context omitted.

[flagged]

I definitely do know what I think.

Do you think developing countries are just peppered with libraries, and their inhabitants order books from Amazon?

Libgen originated in Russia, and its users are global. This is not a purely American issue.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#423
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

" If you plug a laptop into a closet at MIT to download some scientific papers you forfeit your life." This is exactly what I immediately thought while reading the article. It almost feels like the legal system only punishes general public, while most of these guys are above it.

How so? It is still illegal if meta does it, they will face trial.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#424

Earlier quoted context omitted.

It's a stupid situation, though. There are many creators I'm happy to support - but for 99% of them, I don't want their stupid merch . It's mostly low-quality garbage with high markup, that nevertheless cost something to design and produce, thus wasting both precious resources and labor - an useless tax on contributions to artists that doesn't even help anything. I really wish this wasn't necessary. (Even the okay-qu…

I don’t think I’ve ever seen a band selling merch that didn’t also have a tip jar.

I've never seen a band that had. I usually end up buying CDs that then end up on the shelf or in a drawer, never opened.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#425
post #384
post #335

Earlier quoted context omitted.

What are we actually worried about happening? Are AI-written books getting published? If they start out-competing humans, is that bad? According to most naysayers, they can't do anything original. Are people asking the AI for books? And then hoping it will spit it out a human-written book word for word?

> Are AI-written books getting published? Yes, online bookstores are full of them: https://www.nytimes.com/2023/08/05/travel/amazon-guidebooks-... The issue is there's an asymmetry between buyer/seller for books, because a buyer doesn't know the contents until you buy the book. Reviews can help, but not if the reviews are fake/AI generated. In this case, these books are profitable if only a few people buy them as the…

This really has fuck-all to do with copyright though, correct?

If you can't tell how the content is before you read it, it could be written by a monkey.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#426
post #166

Earlier quoted context omitted.

" If you plug a laptop into a closet at MIT to download some scientific papers you forfeit your life." This is exactly what I immediately thought while reading the article. It almost feels like the legal system only punishes general public, while most of these guys are above it.

Airbnb and Uber have showed us that laws matter only to the extent that the political will to enforce them exists. Throw enough lawyers and lobbying money at the problem and the laws can simply be re-written to be friendlier to your business model.

I've also heard the term "regulatory arbitrage" to describe this.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#427

Earlier quoted context omitted.

> Spotify's music library was also pirated in the early days. I want to know more, please enlighten me (anyone who knows). I read the book "The Spotify Play" and it made it seem like the pirated music was an internal-only thing and not something available to customers. Is that true?

Users would upload their copies of the music and spotify would replay them. This was obvious to early users, even if they were only consumers, because of the pirate-shout-out-overlays that were in a lot of the poorer quality releases. Another interesting note, in the early days of spotify, the app would saturate your upload bandwidth while using it. Given their close ties to utorrent, I always assumed that's how they…

Afaik, the trick was to stream (via http, I assume) the first few hundred kilobytes or so from fast/expensive servers and then try to p2p the rest in some clever order. I guess seeking also triggered the fast/expensive path for a while.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#428

Earlier quoted context omitted.

I don't think I've heard the term "English empire". Is it an attempt by the Scottish to pretend they weren't involved?

Is this an attempt to imply the Scots had imperial ambitions and have not been fighting to keep their homes free of invasion for several thousand years? Fuck this sounds familiar right now

>Is this an attempt to imply the Scots had imperial ambitions and have not been fighting to keep their homes free of invasion for several thousand years?

Obviously yes. Who have they been fighting to avoid invasion exactly?

>Fuck this sounds familiar right now

Maybe you read it in a history textbook.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#429
post #335

Earlier quoted context omitted.

What are we actually worried about happening? Are AI-written books getting published? If they start out-competing humans, is that bad? According to most naysayers, they can't do anything original. Are people asking the AI for books? And then hoping it will spit it out a human-written book word for word?

I think the concern goes to the point of copyright to begin with, which is to incentive people to create things. Will the inclusion of copyrighted works in llm training (further) erode that incentive? Maybe, and I think that's a shame if so. But I also don't really think it's the primary threat to the incentive structure in publishing.

> the point of copyright to begin with, which is to incentive people to create things

Is it?

(I don't agree)

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#430
post #281

Earlier quoted context omitted.

Is that a prescriptive or descriptive law?

He left out part of the quote, which is misappropriated as well. Wikipedia: > This quotation is often incorrectly attributed to Francis M. Wilhoit: > Conservatism consists of exactly one proposition, to wit: There must be in-groups whom the law protects but does not bind, alongside out-groups whom the law binds but does not protect. > However, it was actually a 2018 blog response by 59-year-old Ohio composer Frank Wi…

A restatement of Orwell's "all animals are equal, but some are more equal than others".

The irony must have been lost on him.

Post reply on HN