Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

271–280 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#271
post #256

Earlier quoted context omitted.

The reason there was no political will to punish Airbnb and Uber for violating the law was that initially they were subsidized with VC money and so were able to undercut traditional hotels and taxis on price. In the world of tradable goods, pricing below cost with the intent of putting competition out of business so you can raise prices later is known as "dumping" and is itself illegal.

Is it? Really? Or is it just "illegal" for an overseas competitor to a domestic industry, in trade disputes? What is the fine? How many days in jail does the company spend? What portion is its stock diluted by? We remember the tale of Jeff Bezouis the Wise, who tragically lost his company when he decided he didn't want to buy diapers.com at the offered price, and instead wanted to dump 200 million dollars into sellin…

You're right. Dumping refers to international trade. I believe parent commenter was thinking of https://en.wikipedia.org/wiki/Predatory_pricing

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#274

Earlier quoted context omitted.

> Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. I don't understand why you wouldn't just buy copies of the books. Seems like such a relatively inexpensive way to strengthen your legal case.

Buying a copy of the book doesn’t grant you the right to copy it. That is what copyright is for .

They might even have gotten away with legitimate use argument if it was not seeded.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#275
post #60

Earlier quoted context omitted.

First punish them. Then change the laws.

I bet you and my "first build the product, then worry about security" manager would get along.

That one is tough, because they are blind to the risk. I try to only work with people who have been burned before or have been around long enough to have seen the aftermath. Let me guess, they are probably telling you "show me the vulnerability", but refuse to delay shipping or fund the PoC.

Best advice is to communicate in writing the most likely risk and threat scenarios, with as much data or extrapolated data as possible. When the security flaws are later discovered, that is data you can refer to.

From what I read, this is what Zoom was like early on. They had amateur hour security and then when s*t hit the fan they beefed it up and retained a security team. I guess you could say it worked for them?

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#276
post #166

Earlier quoted context omitted.

" If you plug a laptop into a closet at MIT to download some scientific papers you forfeit your life." This is exactly what I immediately thought while reading the article. It almost feels like the legal system only punishes general public, while most of these guys are above it.

Airbnb and Uber have showed us that laws matter only to the extent that the political will to enforce them exists. Throw enough lawyers and lobbying money at the problem and the laws can simply be re-written to be friendlier to your business model.

The hotel and taxi industry were legit terrible before those two disrupted them.

Laws are ment to be broken. Especially in cronist systems where incumbents write the laws.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#277
post #123

Libgen is a civilizational project that should be endorsed, not prosecuted. I hope one day people will look at it and think how stupid we were today to shun the largest collection of literary works in human history.

I think you’re overstating its importance. The internet already makes it possible to order almost any book in existence and have it arrive at your doorstep within a week or so, or often on your ebook reader instantly. And your local library probably participates in an interlibrary loan system that lets you request any book held by any library in the country for free . LibGen gives you access to a much smaller body of…

> LibGen gives you access to a much smaller body of works than either of those.

> Just go to a real library.

The thrill of waiting a week for a book to arrive or navigating the labyrinthine interlibrary loan system is truly a privilege that many can afford. And who needs instant access to knowledge when you can have the pleasure of paying for shipping or commuting to a physical library?

It's also fascinating that you mention compensating authors, as if the current publishing model is a paragon of fairness and equity. I'm sure the authors are just thrilled to receive their meager royalties while the rest of the industry reaps the benefits.

LibGen, on the other hand, is a quaint little website that only offers access to a vast, sprawling library of texts, completely free of charge and accessible to anyone with an internet connection. I'm sure it's totally insignificant compared to the robust and equitable systems you mentioned.

Your suggestion to "just go to a real library" is also a brilliant solution, assuming that everyone has the luxury of living near a well-stocked library, having the time and resources to visit it, and not having any other obligations or responsibilities. I'm sure it's not at all a tone-deaf, out-of-touch recommendation.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#278
post #123

Earlier quoted context omitted.

I think you’re overstating its importance. The internet already makes it possible to order almost any book in existence and have it arrive at your doorstep within a week or so, or often on your ebook reader instantly. And your local library probably participates in an interlibrary loan system that lets you request any book held by any library in the country for free . LibGen gives you access to a much smaller body of…

There are a whole lot of books that are out of print, and if a book went out of print before ebooks were a thing, it probably doesn't have a legal digital edition either.

This. Few people here would remember ebooksclub/gigapedia/smiley/library.nu [1] which predated LibGen by several years. But that online library had a lot of books that are not availble nowadays. There were lots of scanned books (djvu) that people uploaded. So much lost knowledge.

[1] https://en.wikipedia.org/wiki/Library.nu

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#279
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

[flagged]

Shame though that corpos aren't responsible for their actions too :|

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#280
post #186

Earlier quoted context omitted.

Are you sure that everything should be in the public domain? Say you spend a year writing a book, shouldn't you be able to sell it?

They should sell for a price it would make pirating it pointless. Like what Spotify or Netflix did to audio and visual content. Then they can use the exposure to find other ways to make money.

Or if you don't agree with the price, do not buy it. You are NOT entitled to entertain yourself in any way you want. (unless it's funded by taxes etc, in which case... okay, it's open to discussion.)

Look, let's be honest - what gives you or others the right to steal from others?

Post reply on HN