Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

431–440 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#431

Earlier quoted context omitted.

Wilhoit’s law: > There must be in-groups whom the law protects but does not bind, alongside out-groups whom the law binds but does not protect.

Is that a prescriptive or descriptive law?

Lol, that's clearly a descriptive law/maxim not an actual law.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#432

Earlier quoted context omitted.

I rooted for Uber to smash the Taxi cartels. Let us not forget that Taxi Cartels were also insidious beasts. Taxi drivers abused their walled garden with their price gouging by taking longer routes, refusal to take a credit card, and extremely poorly maintained fleets of vehicles. I have had mostly good experiences with Uber, whereas I had experiences that mostly bordered on general condescension toward me whenever I…

What's interesting is that in many cities now, Uber and Lyft are in fact more expensive than taxis. And the experience is equally mediocre. The pendulum has swung back the other way. The only thing they have going for them now is the app based convenience, which is eroding as more "yellow cab" type traditional taxis band together and get set up with their own sort of city-specific app.

> What's interesting is that in many cities now, Uber and Lyft are in fact more expensive than taxis. And the experience is equally mediocre.

That, of course, was the plan all along. Such august figures as JP Morgan, Cornelius Vanderbilt, John D. Rockefeller,and Andrew Carnegie all made their fortune by undercutting the competition, putting them out of business through means legal and otherwise, and finally monopolizing the markets. https://en.wikipedia.org/wiki/Robber_baron_(industrialist)

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#433
post #296

Earlier quoted context omitted.

It grants you the right to read & study it though.

The right to read and study you have by default . It's getting your hands on a book that has legal caveats attached.

Yes, but getting your hands on the material isn't a very interesting legal question IMO.

Whether you can train your LLM on it is a very interesting question.

I've personally never been in favor of punishing people for downloading (or seeding) things.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#434
post #335
post #209

Earlier quoted context omitted.

Libgen turns into a problem when you have a company developing generative AI with it, either giving money to GPU manufacturers or themselves with paid services (see OpenAI)

What are we actually worried about happening? Are AI-written books getting published? If they start out-competing humans, is that bad? According to most naysayers, they can't do anything original. Are people asking the AI for books? And then hoping it will spit it out a human-written book word for word?

>What are we actually worried about happening?

Few company can amass such quantities of knowledge and leverage it all for their own, very-private profits. This is unprecedented centralization of power, for a very select few. Do we actually want that? If not, why not block this until we're sure this a net positive for most people?

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#435
post #329

Earlier quoted context omitted.

Everyone on here is smart enough. Just do not participate and save your money. Do not pay for digital goods. If Netflix raises their prices, it doesn't matter because there is a torrent of all of their shows. If Spotify raises their prices, it doesn't matter because your favorite artist has their entire library in a torrent. If some game company ask you to pay real life prices for a digital costume, find the crack on…

I just can't get behind the sentiment that the unethical behavior by big companies means I get to access all the content I want for free.

They have no morals, therefore I shouldn't either! That'll teach 'em!

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#436
post #391
post #335

Earlier quoted context omitted.

What are we actually worried about happening? Are AI-written books getting published? If they start out-competing humans, is that bad? According to most naysayers, they can't do anything original. Are people asking the AI for books? And then hoping it will spit it out a human-written book word for word?

> Are AI-written books getting published? actually i think they are. lots of e-book slop > If they start out-competing humans, is that bad? Not inherently, but it depends on what you mean by out-competing. Social media outcompeted books and now everyone's addicted and mental illness is more rampant than ever. IMO, a net negative for society. AI books may very well win out through sheer spam but is that good for us?

Nobody has responded to me with anything about how authors are harmed, so I don't really get who we're protecting here.

It feels more like we just want to punish people, particularly rich people, particularly if they get away with stuff we're afraid to try.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#437
post #341

Earlier quoted context omitted.

The traditional taxi industry was rife with corruption, bad experiences, and poor service in many jurisdictions before uber/lyft. As terrible of a human being that I think Travis Kalanick is, it was only going to take lawbreaking to overcome such a tainted system. Medallion systems often prevented any competition, sometimes to absurd effect. The number of licenced taxis often didn't keep pace with population growth,…

Moreover, I believe Uber fundamentally solved two problems with taxis: The driver can't scam the passenger. The driver can't set the meter wrong, drive an unnecessarily long route, or just be an outright unlicensed taxi. Instead, the driver maintains a relationship with Uber, and the passenger can preview the fare before committing. The passenger can't scam the driver. In a traditional taxi, you could theoretically j…

> The driver can't scam the passenger

> The passenger can't scam the driver.

Progress! Uber scams both passenger and driver. Hooray for Free Markets™!

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#438

Earlier quoted context omitted.

The traditional taxi industry was rife with corruption, bad experiences, and poor service in many jurisdictions before uber/lyft. As terrible of a human being that I think Travis Kalanick is, it was only going to take lawbreaking to overcome such a tainted system. Medallion systems often prevented any competition, sometimes to absurd effect. The number of licenced taxis often didn't keep pace with population growth,…

in my experience, taxi quality varies wildly depending on where you are in bay area, it absolutely makes sense to invent uber, because the taxis were awful. and in vancouver (canada), they're also awful, and deserve the disruption: they would often tell you it'd be a 40 minute wait, and then just not show up taxis in new york were and continue to be totally fine. you just stand outside and get in ~20 seconds later, w…

>taxis in new york were and continue to be totally fine. you just stand outside and get in ~20 seconds later, with no hassles or apps

Unless you weren't white, or you wanted to leave Manhattan (or even go north of 96th street). Otherwise, yeah I guess they were okay.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#439
post #335

Earlier quoted context omitted.

What are we actually worried about happening? Are AI-written books getting published? If they start out-competing humans, is that bad? According to most naysayers, they can't do anything original. Are people asking the AI for books? And then hoping it will spit it out a human-written book word for word?

>What are we actually worried about happening? Few company can amass such quantities of knowledge and leverage it all for their own, very-private profits. This is unprecedented centralization of power, for a very select few. Do we actually want that? If not, why not block this until we're sure this a net positive for most people?

Meta open-sourced it my guy

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#440
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

Yes. And the problem here isn't that companies get away with doing things like this, the problem is that individuals don't. Attempting to lock information behind a nightmarish legal system is the problem. I'm pretty much at the point now where I don't buy the "copyright incentivizes creation" argument any more. Copyright, like advertising, incentivizes creation by enormous corporations, but also like advertising it i…

Sure creative people will always create but the scope of that creativity will be limited if we do away with intellectual property. Steve Spielberg would probably always have created movies, but he wouldn't have been able to make Jurassic Park, Saving Private Ryan,or Indiana Jones without capital from the studio system, and the studio system wouldn't have provided him with that capital of they couldn't extract economic rents from the copyright for those films.
Post reply on HN