Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

451–460 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#451
post #193
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

i know of a company that poisoned an entire town! thats terrorism if done by an individual. the company still exists, just paid a settlement and carried on...

Are you talking about Bhopal in 1984? If so it would be an understatement to refer to half a million people as a “town”, and an overstatement to imply it was terrorism. Willful negligence, yes, but terrorism, no.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#452

Earlier quoted context omitted.

The traditional taxi industry was rife with corruption, bad experiences, and poor service in many jurisdictions before uber/lyft. As terrible of a human being that I think Travis Kalanick is, it was only going to take lawbreaking to overcome such a tainted system. Medallion systems often prevented any competition, sometimes to absurd effect. The number of licenced taxis often didn't keep pace with population growth,…

It sucked, but not everywhere equally. Meanwhile, Uber rode their one-trick pony (an app), which everyone quickly replicated, all the way to upending taxi businesses worldwide , thanks to their infinite money supply letting them survive long enough in any new market to get the public behind them, which took away support from local regulators trying to keep the market from being gutted by what at this point was a mult…

That seems a little dramatic. They never forced anyone to take an uber right? If taxis were so amazing in other countries why would anyone be interested in switching to uber?

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#453
post #341

Earlier quoted context omitted.

The traditional taxi industry was rife with corruption, bad experiences, and poor service in many jurisdictions before uber/lyft. As terrible of a human being that I think Travis Kalanick is, it was only going to take lawbreaking to overcome such a tainted system. Medallion systems often prevented any competition, sometimes to absurd effect. The number of licenced taxis often didn't keep pace with population growth,…

Moreover, I believe Uber fundamentally solved two problems with taxis: The driver can't scam the passenger. The driver can't set the meter wrong, drive an unnecessarily long route, or just be an outright unlicensed taxi. Instead, the driver maintains a relationship with Uber, and the passenger can preview the fare before committing. The passenger can't scam the driver. In a traditional taxi, you could theoretically j…

Meanwhile, in places with sensible rules about taxis and private hire, the only thing that Uber did was make it easier for people to break the rules. And rack up an enormous tax bill that they somehow believed they'd be able to get out of paying.

https://www.londonreconnections.com/2021/uber-loses-appeal-a...

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#454

I strongly urge people to read Thomas Babington Macaulay's speeches on copyright, its aims, terms, and hazards. Very well reasoned and explained. In particular, people often cited the case of authors who had died leaving a family in destitution, and claimed that copyright extension would be a fair way of preventing this, but in most cases the remaining family had never held the copyright; the author had initally sold…

This one example does not make stealing acceptable which is what you’re implying.

Copyright infringement isn't stealing. I will die on this hill!

also, I don't think that implication is required, but lets pretend the implication is the only reasonable conclusion one could draw. Maybe it does make it acceptable?

If the vast majority of copyright enforcement isn't to protect creators of valuable work, but only serves to enrich those who take advantage of those creators. Then isn't it not just reasonable or acceptable, but ethically required for someone to do everything they can to dismantle the systems they're abusing against the interests of those who are actually improving the world with their creations?

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#455
post #41

Considering prices for single work, this must be multi-billion dollar compensation. Take for example 675k paid for 31 songs. So 20k a song. If we estimate book to be say 10MB that would 8 million works. So I think reasonable compensation is something along 163 billion. Not even 10 years of net income. Which I think is entirely fair punishment.

Meta argues that it's fair use, and that they just downloaded, and never seeded, all the torrents.

Seeding and downloading are in the same protocol. You can't do one without the other

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#456
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

" If you plug a laptop into a closet at MIT to download some scientific papers you forfeit your life." This is exactly what I immediately thought while reading the article. It almost feels like the legal system only punishes general public, while most of these guys are above it.

It's not "almost" like that. The legal system IS that.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#457
post #425
post #384

Earlier quoted context omitted.

> Are AI-written books getting published? Yes, online bookstores are full of them: https://www.nytimes.com/2023/08/05/travel/amazon-guidebooks-... The issue is there's an asymmetry between buyer/seller for books, because a buyer doesn't know the contents until you buy the book. Reviews can help, but not if the reviews are fake/AI generated. In this case, these books are profitable if only a few people buy them as the…

This really has fuck-all to do with copyright though, correct? If you can't tell how the content is before you read it, it could be written by a monkey.

This is starting to get pretty circular. The AI was trained on copyrighted data, so we can make a hypothesis that it would not exist - or would exist in a diminished state - without the copyright infringement. Now, the AI is being used to flood AI bookstores with cheaply produced books, many of which are bad, but are still competing against human authors.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#458
post #117

Earlier quoted context omitted.

Parent comment implies Swartz was guilty of some degree. I vehemently disagree with that.

> Parent comment implies Swartz was guilty of some degree as a constructive criticism, you might want to reconsider your interpretation of >"Remembering Aaron Swartz in this moment" -> Which was arguably more innocent — scientific papers. As in, both hold some degree of illegality (objectively), so when pointed that "he is guilty of some degree" is due to the jurisdiction laws (broken or not) regardless of societal/m…

You're quoting a different user

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#459

Earlier quoted context omitted.

The reason there was no political will to punish Airbnb and Uber for violating the law was that initially they were subsidized with VC money and so were able to undercut traditional hotels and taxis on price. In the world of tradable goods, pricing below cost with the intent of putting competition out of business so you can raise prices later is known as "dumping" and is itself illegal.

I rooted for Uber to smash the Taxi cartels. Let us not forget that Taxi Cartels were also insidious beasts. Taxi drivers abused their walled garden with their price gouging by taking longer routes, refusal to take a credit card, and extremely poorly maintained fleets of vehicles. I have had mostly good experiences with Uber, whereas I had experiences that mostly bordered on general condescension toward me whenever I…

It is possible that digitization and improvement of taxi services was inevitable anyway

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#460
The only bad thing about this is that small time players who do it are treated poorly (Aaron Swartz). IP de-facto not existing for AI companies is a feature, not a bug.

The fact that most of the world embraced hardcore copyright troll ludditism when the means of their (badly paying creative) jobs economic production was democratized implies that most people do not believe in any "egalitarianism" and especially not the left-wing form many profess to believe in. Certainly not "information wants to be free" or any of the other idealist shit that I or Aaron Swartz believed in. What meta did was software communism - full stop. They literally released their models to the public! I support all of this 10000%. The only issue is that they're not open enough (fully open source the dataset)

So, unironically, good! Thank you, please pirate more! Please destroy the US IP system while you're at it. Copyright abolitionism is good and thank you Zuckerberg!

Post reply on HN