Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

621–630 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#621
Wouldn't it be a real shame if the entirety of US constitution, laws, and legal precedence went out the window these days, and the only thing left unscathed was the rotten mess that is copyright law? Just saying, this might be the moment to burn it to the ground. Not that it makes up for any of the other stuff going on, but why waste a perfectly good crisis?

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#622

Earlier quoted context omitted.

Do you have a citation for that claim? I've not seen a claim that none of the material had copyright before.

It's a library of historical scientific work. You will find the famous Einstein's 3 1905 papers there, for example.

Every scientific paper in the last 90 years or so is still under copyright, owned by the authors, the published, or the universities.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#623

Earlier quoted context omitted.

This argument ignores the fact that there were other alternatives to Uber at the time, ones that didn't break the law! Believe it or not, there were multiple ride hailing apps on the iPhone, but none were as great at accumulating capital or breaking the law without recourse.

The laws they allegedly broke were the taxi medallion cartel laws, which were the things keeping taxis terrible by limiting supply and competition. And those laws in general apply to the drivers rather than the ride hailing service. There is also a lot of ambiguity there, e.g. if you have a ride sharing service where people go on the app to find people to carpool with on a trip they'd be making anyway and then contri…

Taxi Medallion laws were also a Reputation Engine that was publicly queryable, subject to FOIA laws and generally had easy to search public databases for them, with detailed notes. Sure Uber/Lyft boil that into a "friendly" 5-star UI, but do you have any idea what data contributed to that star rating? Do you always trust the algorithms that compute them from a bucket of metrics you can't directly request?

Sure, Medallion laws had problems, and got Regulatory Captured in some cities to also become terrible Trusts controlling prices that needed busting. But the answer to "fix the Regulation" isn't always "break the Regulation", and the Regulation had a lot of good intent of having public accessible information about drivers and that data not just owned by a single company and locked in their opaque algorithms. It might have been nicer to fix the Regulatory Capture and Bust the Trusts.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#624
post #498

Earlier quoted context omitted.

VC and startups are fundamentally about disruption. You can't make an omelette without breaking a few eggs (laws). The incumbent players are not going to sit still and let things be "disrupted". A common response is to make sure the public knows about the broken eggs. I would say youtube, Google, Spotify, Uber, doordash, etc. all have made my life much better.

You don't know a world without them so you actually have no idea if they have made your life compared to that world much better or much worse. How your life was at the time is irrelevant.

This is a vacuous statement. You can say the same thing about electricity, or antibiotics, or any other modern advancement.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#625

Earlier quoted context omitted.

The purpose of having an executive branch of government is explicitly to apply the law based on subjective opinions. There's no purpose of having an executive branch of government separate from the other two branches if not to cushion the inflexible and glacial nature of the other branches of government.

>The purpose of having an executive branch of government is explicitly to apply the law based on subjective opinions. What? No, the purpose of having a separate executive is separation of powers and checks and balances.

You haven't explained why there is an executive branch in the first place.

Why does the executive branch exist at all if it's simply to enforce written law?

Why do we elect the executive at all if they are merely to enforce written law?

Why do executives have the power to pardon someone when a court of law finds a person guilty of breaking law?

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#626
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

Some can pirate on a large scale and see no repercussions. Some can steal from stores and see no repercussions. Some can steal from others and see no repercussions. Some can violently harm others and see no repercussions. Some can damage property and see no repercussions. Some can’t. This world is not right.

The strong do what they wish and the weak suffer what they must. Any morality beyond that is a fairy tale that the weak tell themselves.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#627
post #77

This should be legal. Copyright law does more harm than good. The only ethical problem here is that only Meta sized companies can afford to pay the "damages" for such blatant law violations at worst, or the fees of their lawyers at best.

If an individual was the one tormenting almost 82 TB of copyrighted books, the damages they would have to pay would be in the trillions (mostly because of how broken the copyright law system is)

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#628
post #609

Earlier quoted context omitted.

Everyone on here is smart enough. Just do not participate and save your money. Do not pay for digital goods. If Netflix raises their prices, it doesn't matter because there is a torrent of all of their shows. If Spotify raises their prices, it doesn't matter because your favorite artist has their entire library in a torrent. If some game company ask you to pay real life prices for a digital costume, find the crack on…

lol I absolutely do not want non digital goods nor pirating. Ever. It's 2025. I don't have a cdplayer, a tape player, a blue ray player, I don't even know what the most modern "blue ray" disc would be. I have $2k worth of vinyls that are just unique copies I display as art I'll never put in my record player, that's also never been used. I don't want to constantly worry about 60gb of mp3 files. Oh no, that TV show I'l…

>MSNBC just cancelled Andrea Mitchells TV show, today, because she brought in no younger audiences. So yes, shows do get cancelled by not being watched.

Did anyone, young or old, want to watch an 80 year old stumble over her words, lose her train of thought, and speak so painfully slow? She had built up connections over her long career but was basically unwatchable. The worst part of a Kamala presidency would have been her on the news and not in retirement.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#629
post #609

Earlier quoted context omitted.

lol I absolutely do not want non digital goods nor pirating. Ever. It's 2025. I don't have a cdplayer, a tape player, a blue ray player, I don't even know what the most modern "blue ray" disc would be. I have $2k worth of vinyls that are just unique copies I display as art I'll never put in my record player, that's also never been used. I don't want to constantly worry about 60gb of mp3 files. Oh no, that TV show I'l…

>MSNBC just cancelled Andrea Mitchells TV show, today, because she brought in no younger audiences. So yes, shows do get cancelled by not being watched. Did anyone, young or old, want to watch an 80 year old stumble over her words, lose her train of thought, and speak so painfully slow? She had built up connections over her long career but was basically unwatchable. The worst part of a Kamala presidency would have be…

[deleted]

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#630
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

Reminds me of recent discussions about similar topic, what may clearly look like a crime can be treated differently depending on if you do it as an individual or as a company. Somewhere down the line its all about understanding the limits and boundaries of the system, its a skill in itself.
Post reply on HN