Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

471–480 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#471

I strongly urge people to read Thomas Babington Macaulay's speeches on copyright, its aims, terms, and hazards. Very well reasoned and explained. In particular, people often cited the case of authors who had died leaving a family in destitution, and claimed that copyright extension would be a fair way of preventing this, but in most cases the remaining family had never held the copyright; the author had initally sold…

This one example does not make stealing acceptable which is what you’re implying.

[deleted]

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#472
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

I think the difference may be LLMs may not be laundered clean of copyright data anytime soon. Even if chatgpt got big and profitable, it's not so clear that it won't contain copyrighted data as that may simply be necessary to train the best models.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#473

Earlier quoted context omitted.

I agree with your point, but will split hairs on using the word "terrorism". I think that should be reserved for people that commit atrocities for some political aim. I'm fairly sure the company in question (I assume Union Carbide) did not poison the town to advance a political agenda.

Kinda terrifying that you can get away with shit if you just argue that it's not politically motivated, you did it because you really wanted a yacht.

You also get a lesser offense killing someone accidentally as opposed to a premeditated murder.

There’s a difference between an intentional act, an accident, and an accident due to extreme neglect and our laws reflect that.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#474
post #341

Earlier quoted context omitted.

Moreover, I believe Uber fundamentally solved two problems with taxis: The driver can't scam the passenger. The driver can't set the meter wrong, drive an unnecessarily long route, or just be an outright unlicensed taxi. Instead, the driver maintains a relationship with Uber, and the passenger can preview the fare before committing. The passenger can't scam the driver. In a traditional taxi, you could theoretically j…

> The driver can't scam the passenger > The passenger can't scam the driver. Progress! Uber scams both passenger and driver. Hooray for Free Markets™!

Under what definition of “scam” is that the case?

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#475
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

> Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it

To this day, there are a huge number of videos that show copyrighted content on YouTube; they are usually crappy clips, reversed and with different music playing in the background to avoid automated detection.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#476
Is there a concept in the legal system of first-come-first-served that could be used as precedent?

What I mean is: when someone is prosecuted for copyright infringement, but Meta isn't, then could the case be put on hold until Meta is found guilty and pays a fine?

Also maybe the fine on the later case would have to be proportional to the prior case. So if Meta pays $1 per infringement, the penalty might be $1 for torrenting something else (which is immaterial and not worth the justice system's time) so pretty much all copyright infringement cases would get thrown out.

It reminds me of how mainstream drug addicts get convicted and spend years in prison, while celebrities get off with a warning or monetary fine.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#477
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

So be a company? Last I checked it costs a couple of hundred dollars to form an LLC, what am I missing?

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#478
post #341

Earlier quoted context omitted.

Moreover, I believe Uber fundamentally solved two problems with taxis: The driver can't scam the passenger. The driver can't set the meter wrong, drive an unnecessarily long route, or just be an outright unlicensed taxi. Instead, the driver maintains a relationship with Uber, and the passenger can preview the fare before committing. The passenger can't scam the driver. In a traditional taxi, you could theoretically j…

> The driver can't scam the passenger > The passenger can't scam the driver. Progress! Uber scams both passenger and driver. Hooray for Free Markets™!

> Uber scams both passenger and driver.

Then why do people keep using it? It seems like a pretty transactional relationship to me. If drivers aren't getting paid as much as they wanted, they should find another job with a higher price. If passengers are paying more than they wanted, then they should find another way to call a taxi with a lower price.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#479
post #329

Earlier quoted context omitted.

I just can't get behind the sentiment that the unethical behavior by big companies means I get to access all the content I want for free.

They have no morals, therefore I shouldn't either! That'll teach 'em!

A whole lot of people in the tech scene got really mad when Huawei was using obviously stolen Cisco designs and code for their switches. Didn’t humanity benefit from having cheaper access to switches because they didn’t have to pay for Cisco’s sunk costs? A whole lot of people got mad when Microsoft reportedly ganked open source code for things like DNS. Didn’t humanity benefit from one of the world’s most popular server OSs having more reliable name resolution?

Oh, but corporations were the primary beneficiaries, right?

Well, corporations are the primary beneficiaries of this too from a financial perspective. A vanishingly small percentage of people will run, let alone train these models themselves— it’s almost exclusively used to make commercial services that directly compete against the people that made the initial ” data“. But, the vanishingly small percentage of people that directly utilize this stuff for non-commercial use frequent echo chambers like this that make them think more regular people benefit directly. And the companies that are competing directly with creatives and intellectuals using their stolen work employ a whole lot of people here, directly or not.

The distinction between a reason and a justification gets pretty difficult to distinguish the closer you are to the group benefiting from injustice.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#480

Earlier quoted context omitted.

The traditional taxi industry was rife with corruption, bad experiences, and poor service in many jurisdictions before uber/lyft. As terrible of a human being that I think Travis Kalanick is, it was only going to take lawbreaking to overcome such a tainted system. Medallion systems often prevented any competition, sometimes to absurd effect. The number of licenced taxis often didn't keep pace with population growth,…

in my experience, taxi quality varies wildly depending on where you are in bay area, it absolutely makes sense to invent uber, because the taxis were awful. and in vancouver (canada), they're also awful, and deserve the disruption: they would often tell you it'd be a 40 minute wait, and then just not show up taxis in new york were and continue to be totally fine. you just stand outside and get in ~20 seconds later, w…

> taxis in new york were and continue to be totally fine. you just stand outside and get in ~20 seconds later, with no hassles or apps.

This is only true in a small subset of New York.

Post reply on HN