Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

341–350 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#341

Earlier quoted context omitted.

The reason there was no political will to punish Airbnb and Uber for violating the law was that initially they were subsidized with VC money and so were able to undercut traditional hotels and taxis on price. In the world of tradable goods, pricing below cost with the intent of putting competition out of business so you can raise prices later is known as "dumping" and is itself illegal.

The traditional taxi industry was rife with corruption, bad experiences, and poor service in many jurisdictions before uber/lyft. As terrible of a human being that I think Travis Kalanick is, it was only going to take lawbreaking to overcome such a tainted system. Medallion systems often prevented any competition, sometimes to absurd effect. The number of licenced taxis often didn't keep pace with population growth,…

Moreover, I believe Uber fundamentally solved two problems with taxis:

The driver can't scam the passenger. The driver can't set the meter wrong, drive an unnecessarily long route, or just be an outright unlicensed taxi. Instead, the driver maintains a relationship with Uber, and the passenger can preview the fare before committing.

The passenger can't scam the driver. In a traditional taxi, you could theoretically just walk out ("dine and dash" style). The passenger can also make a call to dispatch and not show up for the ride. Instead, the passenger maintains a relationship with Uber, and the driver doesn't need to handle any payments.

> Medallion systems often prevented any competition, sometimes to absurd effect. The number of licenced taxis often didn't keep pace with population growth, sometimes even staying flat.

And thus medallion owners collect economic rent on their artificially scarce resource, distorting the free market. https://en.wikipedia.org/wiki/Economic_rent

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#342

Earlier quoted context omitted.

The hotel and taxi industry were legit terrible before those two disrupted them. Laws are ment to be broken. Especially in cronist systems where incumbents write the laws.

Hotels were just fine. Taxis were discriminatory and "uncool" to the point were Uber has saved thousands by preventing drunk driving. Now if you go out with the boys and get drunk, it's a 30 second casual call to get an Uber and get home. Live in a neighborhood Taxis are afraid to service,you can either make some extra income working for Uber or use it yourself. When Ubers used as its intended purpose, to basically m…

> Say your rents it's going to be late, you can pick up 20 or 30 hours of Uber this month to make it happen

that sounds so incredibly dystopian, not sure if that was the intention :(

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#343
post #123

Earlier quoted context omitted.

I think you’re overstating its importance. The internet already makes it possible to order almost any book in existence and have it arrive at your doorstep within a week or so, or often on your ebook reader instantly. And your local library probably participates in an interlibrary loan system that lets you request any book held by any library in the country for free . LibGen gives you access to a much smaller body of…

> LibGen gives you access to a much smaller body of works than either of those. > Just go to a real library. The thrill of waiting a week for a book to arrive or navigating the labyrinthine interlibrary loan system is truly a privilege that many can afford. And who needs instant access to knowledge when you can have the pleasure of paying for shipping or commuting to a physical library? It's also fascinating that you…

Your library almost definitely offers digital loans as well.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#345
post #185
post #58

We all like hating big corporations, especially Meta, and people seem to use this as an opportunity to advocate for punishing them. I think it's wiser to advocate for changing our IP laws.

You're conflating different problems. Big corporations are too big, they should just not exist. When you have corporations more powerful than the government of the biggest states, it's a bug, not a feature. The IP laws may need rethinking. Saying that they should disappear because big corporations are above the law doesn't help, though. First kill the big corporations, then think about fair laws. Changing the law now…

How do you suggest making them smaller?

For instance, what if google was still just serving search results w/ ads, and they never expanded that. How would you make them smaller?

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#346
post #308

Earlier quoted context omitted.

The traditional taxi industry was rife with corruption, bad experiences, and poor service in many jurisdictions before uber/lyft. As terrible of a human being that I think Travis Kalanick is, it was only going to take lawbreaking to overcome such a tainted system. Medallion systems often prevented any competition, sometimes to absurd effect. The number of licenced taxis often didn't keep pace with population growth,…

I've never been a huge user of either, but my worst Uber ride was much better than my best taxi ride.

The last time I dragged my family into a taxi because of my anti Uber ideology, the driver stank to hell of body odor, asked me to input directions on his phone covered with dried snot from him sneezing with his mouth open, he drove dangerously under the speed limit on the freeway, and it took twice as long to get home as normal.

But at least I didn’t give Uber any money…

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#347
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

" If you plug a laptop into a closet at MIT to download some scientific papers you forfeit your life." This is exactly what I immediately thought while reading the article. It almost feels like the legal system only punishes general public, while most of these guys are above it.

They may have just been the friendly step A. We didn't end up seeing where that was going to go.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#348
post #123

Libgen is a civilizational project that should be endorsed, not prosecuted. I hope one day people will look at it and think how stupid we were today to shun the largest collection of literary works in human history.

I think you’re overstating its importance. The internet already makes it possible to order almost any book in existence and have it arrive at your doorstep within a week or so, or often on your ebook reader instantly. And your local library probably participates in an interlibrary loan system that lets you request any book held by any library in the country for free . LibGen gives you access to a much smaller body of…

Libraries can burn down (see Library of Alexandria), civilizations end (see various). LibGen makes it possible for an individual to backup a snapshot of cumulative human knowledge, and I think that's commendable.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#349
post #185
post #58

We all like hating big corporations, especially Meta, and people seem to use this as an opportunity to advocate for punishing them. I think it's wiser to advocate for changing our IP laws.

You're conflating different problems. Big corporations are too big, they should just not exist. When you have corporations more powerful than the government of the biggest states, it's a bug, not a feature. The IP laws may need rethinking. Saying that they should disappear because big corporations are above the law doesn't help, though. First kill the big corporations, then think about fair laws. Changing the law now…

> First kill the big corporations, then think about fair laws.

It's not possible to kill big corporations before fair laws, because as you said yourself "corporations are already above the law"

Unfair laws don't apply to big corporations, they only apply to the people opposed to big corporations

It's akin to hamstringing a horse and saying you'll fix it when they win

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#350
post #43
post #21

Earlier quoted context omitted.

Meta already does that to themselves every year or so, deleting all internal communications. They've thrown away a huge amount of communication to source code commit reinforcement training data as a result. They do it to avoid emails making it into trials like this.

No large company will ever consider training a public LLM on all their internal communications.

Could be a private finetune, or even a complete private model. They already have one for their internal codebase.
Post reply on HN