Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

411–420 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#411

Earlier quoted context omitted.

I rooted for Uber to smash the Taxi cartels. Let us not forget that Taxi Cartels were also insidious beasts. Taxi drivers abused their walled garden with their price gouging by taking longer routes, refusal to take a credit card, and extremely poorly maintained fleets of vehicles. I have had mostly good experiences with Uber, whereas I had experiences that mostly bordered on general condescension toward me whenever I…

What's interesting is that in many cities now, Uber and Lyft are in fact more expensive than taxis. And the experience is equally mediocre. The pendulum has swung back the other way. The only thing they have going for them now is the app based convenience, which is eroding as more "yellow cab" type traditional taxis band together and get set up with their own sort of city-specific app.

The price people are willing to pay sets how nice a cab fleet can be while still turning a profit.

Same reason you don't see landscaping crews filled out with stellar employees.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#412
post #335
post #209

Earlier quoted context omitted.

Libgen turns into a problem when you have a company developing generative AI with it, either giving money to GPU manufacturers or themselves with paid services (see OpenAI)

What are we actually worried about happening? Are AI-written books getting published? If they start out-competing humans, is that bad? According to most naysayers, they can't do anything original. Are people asking the AI for books? And then hoping it will spit it out a human-written book word for word?

I think the concern goes to the point of copyright to begin with, which is to incentive people to create things. Will the inclusion of copyrighted works in llm training (further) erode that incentive? Maybe, and I think that's a shame if so. But I also don't really think it's the primary threat to the incentive structure in publishing.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#414
post #251
post #11

So according to some AI, the damages awarded per infringed work is ~$750 minimum in the US. 80TB of books, each let's say 10MB on average, would be 8 million works. So Meta should pay 6 billion USD for their copyright infringement?

Prosecutors filed for Swartz 50 years of imprisonment and $1 million in fines. Can you calculate how many years that would be for Mark and his people?

I ran it, it came out to zero

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#415

Earlier quoted context omitted.

The reason there was no political will to punish Airbnb and Uber for violating the law was that initially they were subsidized with VC money and so were able to undercut traditional hotels and taxis on price. In the world of tradable goods, pricing below cost with the intent of putting competition out of business so you can raise prices later is known as "dumping" and is itself illegal.

The traditional taxi industry was rife with corruption, bad experiences, and poor service in many jurisdictions before uber/lyft. As terrible of a human being that I think Travis Kalanick is, it was only going to take lawbreaking to overcome such a tainted system. Medallion systems often prevented any competition, sometimes to absurd effect. The number of licenced taxis often didn't keep pace with population growth,…

in my experience, taxi quality varies wildly depending on where you are

in bay area, it absolutely makes sense to invent uber, because the taxis were awful. and in vancouver (canada), they're also awful, and deserve the disruption: they would often tell you it'd be a 40 minute wait, and then just not show up

taxis in new york were and continue to be totally fine. you just stand outside and get in ~20 seconds later, with no hassles or apps. i've been in an uber/lyft a handful of times in nyc, but they're just worse (possibly cheaper, but the subway also gives them stiff competition, and i don't care that much if i'm in enough of a hurry to take a cab)

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#416

Earlier quoted context omitted.

The hotel and taxi industry were legit terrible before those two disrupted them. Laws are ment to be broken. Especially in cronist systems where incumbents write the laws.

Hotels were just fine. Taxis were discriminatory and "uncool" to the point were Uber has saved thousands by preventing drunk driving. Now if you go out with the boys and get drunk, it's a 30 second casual call to get an Uber and get home. Live in a neighborhood Taxis are afraid to service,you can either make some extra income working for Uber or use it yourself. When Ubers used as its intended purpose, to basically m…

> Say your rents it's going to be late, you can pick up 20 or 30 hours of Uber this month to make it happen.

Maybe... I really don't get how the economics work out here though. If you look at the numbers, it mostly just seems like you're converting car equity into cash via depreciation.

But also, I'd guess that for a big chunk of people who are going to have trouble paying rent with any regularity, they'd have to overpay for their car in the first place to get something that's Uber-appropriate. My car's a couple years too old for Uber now, but is still perfectly functional, and there's just no way the math would work for me to buy a newer car so that I can convert its capital cost into cash via Uber.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#417
post #354

Earlier quoted context omitted.

" If you plug a laptop into a closet at MIT to download some scientific papers you forfeit your life." This is exactly what I immediately thought while reading the article. It almost feels like the legal system only punishes general public, while most of these guys are above it.

> the legal system only punishes general public, while most of these guys are above it It’s because the legal system is not about justice, it’s about money Most people can’t afford lawyers or expensive legal battles On the other hand, individuals and organizations with a lot of money get to weaponize and exploit the legal system to their advantage “To my friends, anything; to my enemies, the law”

At the risk of wading into politics - consider a legal environment, in any country, where laws become increasingly strict, but where prosecutorial discretion, pardon powers, and a justice system designed to allow well-resourced law firms to delay cases indefinitely, are all transparently used for political purposes. Such an environment could easily exhibit a feedback loop that allows justice to be arbitrary and opposition voices to be silenced.

I'll refrain from value judgments on the above - but for heaven's sake, we're on a site called "Hacker News." We should understand that a machine like this could turn on any one of us in an instant for any reason.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#418

Earlier quoted context omitted.

It's a stupid situation, though. There are many creators I'm happy to support - but for 99% of them, I don't want their stupid merch . It's mostly low-quality garbage with high markup, that nevertheless cost something to design and produce, thus wasting both precious resources and labor - an useless tax on contributions to artists that doesn't even help anything. I really wish this wasn't necessary. (Even the okay-qu…

What is the point of this comment? Just a stream of consciousness for a future LLM sweep? Nobody thinks that the actual creator should get nothing. Are you asking for better T shirts? Do you want more direct ways of just giving cash?

The latter, yes.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#419
post #170

Maybe you should go after the worst offender (OpenAI) first before going after Meta, since the latter already gave back their model away for free for everyone and the architecture. We will know why OpenAI isn't getting investigated.

Could be why OpenAI paid them so much, to go after their open-source competition hardest of all.
Post reply on HN