Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

511–520 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#511
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

This frames Google's indexing of the web in a totally, abjectly wrong fashion. It wasn't "other people's data", it was data people published to the public internet, implicitly and explicitly granting permission to download through the act of serving that data without restriction to whoever navigated to a particular URL.

That's how the internet works. If you want private content, you need to put up a gate mechanism of some sort with authentication or other methods of restricting access. Without that, you are literally having your server "serve" the content to whoever asks for it, without restriction or exception, without ToS or meaningful contract or agreements.

You can't have it both ways. "But they didn't know" or other post-hoc claims of innocent people publishing content to the web being misled or confused or abused is infantilizing nonsense.

The web wouldn't have been as amazing and revolutionary and liberating if the fundamental public and open nature of its systems was private and walled off by default.

Your take on YouTube going viral initially over copyrighted content isn't correct, either - it was ease of use and access. It was fairly popular by the time Google bought it, and once it was reachable and advertised by google itself, it exploded, because by that time, everyone had defaulted to using google for search.

Other people corrected your Spotify take.

The reason they pirated is because it is functionally impossible to gain access to the data in any other way. For consumers, there are lots of old shows, music, and other content that aren't accessible, so they turn to piracy. A vast majority of the time, if content is accessible, people will pay and do the technically legal and "right" thing.

Publishers exploit authors and content creators in the name of "platforming" and "marketing" , effectively doing as little as possible to take 90%+ of the value of a product and providing as little as possible to the producer of content or books or music. They get by on technicalities and have captured the legal arena entirely, with any attempt at reform or revolution meeting a messy death at the hands of lawyers and big money publishers.

Screw those people. They lie, cheat, and steal, and somehow have gotten away with fooling the world into thinking they're the good guys.

Copying bits and bytes is not stealing, and the ones trying to shill that narrative are trying to fool as many people as possible into giving them more money without any return of value in kind. I'd download the hell out of a car. Pirate everything.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#512

Earlier quoted context omitted.

I find this such a strange remark on this front. You got less than 1% of a book... from an author who has passed away... who wrote on a research topic that was funded by an institution that takes in hundreds of millions of dollars in federal grants each year... I'm not an author (although I do generate almost exclusively IP for a living) and I think this is about as weak a form of this argument as you possibly make.…

It isn't that someone was hurt. We have one private entity gaining power by centralizing knowledge (which they never contributed to) and making people pay for regurgitating the distilled knowledge, for profit. Few entities can do that (I can't). Most people are forced to work for companies that sell their work to the higher bidder (which are the very entities mentioned above), or ask them to use AI (under the conditi…

Are you talking about Meta? They released the model. It's free to use.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#514

Earlier quoted context omitted.

This is the corollary of the fallacy of appeal to authority: the rejection of an argument on the grounds that the speaker was horribly wrong on an unrelated or very loosely related topic. If you reject Macaulay on copyright because he was an imperialist, you can use the exact same logic to reject the arguments of essentially every person who ever lived. Very few humans who ever wrote anything important will perfectly…

> If you reject Macaulay on copyright because he was an imperialist On the contrary I would argue that this is precisely why you SHOULD NOT take his opinion on copyright. One of the main outcomes of imperialism/colonization is denigrating/destroying/appropriating works of art, literature with the primary goal of subjugation, subversion and thereupon replacement of native culture/traditions/institutions. I did not quo…

> One of the main outcomes of imperialism/colonization is denigrating/destroying/appropriating works of art, literature with the primary goal of subjugation, subversion and thereupon replacement of native culture/traditions/institutions.

Which is irrelevant to the question of whether copyright law within the country of England and within English culture is beneficial or not.

It is the nature of racism that it bypasses rational thought—it does not follow that because someone is racist they therefore don't have anything valuable to say on loosely related topics. Someone can see clearly about copyright when thinking about English authors while treating non-English authors as strictly inferior.

These kinds of contradictions are to be expected when racism is involved, because racism inherently lives in the lizard brain (occasionally justified by post hoc rationalizations). Someone's arguments about an issue touching only their own tribe will tend to be more rational than those that touch on other tribes, and you'll miss out if you assume the rationality is going to be correlated and dismiss all arguments accordingly.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#515
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

> If you plug a laptop into a closet at MIT to download some scientific papers you forfeit your life. In case anybody here doesn't know, that's a reference to Aaron Swartz, an activist (and Reddit co-founder) that was risking 35 years in prison and a $1 million fine just for downloading a lot of academic papers from JSTOR. He eventually took his life because of the pressure. May his soul rest in peace.

Except he was offered 6 months in a plea bargain, which he declined because he wanted a trial. Whether 6 months was reasonable punishment for "plug a laptop into a closet at MIT to download some scientific papers" is another matter, but "you forfeit your life" or "35 years in prison and a $1 million fine " is massively misleading.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#516
For some misterious reason I can't see Zuckerberg in front of a judge facing 50 years imprisonment. Anyone can?

I truly hope that whoever takes the case goes after Meta with 1000 times the pressure that was put on Swartz, but honestly I don't expect much just as the top comment precisly expressed.

And if we are going to be fair please also let's not forget about the other usual suspects, or anyone thinks they are falling behind?

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#518
post #215

Earlier quoted context omitted.

We don’t allow indiscriminate human experimentation in medicine. We have crimes against this, and yet we still have new medicines. Sure, it won’t be as quick if we could just use humans as test subjects from the start, but that’s an unethical line. Innovation done immorally is progress that shouldn’t have been made. The ends don’t justify the means, but I’m not an ethical nihilist. The crime is downloading and copyin…

Those medical policies have condemned thousands, possibly millions, to lives of unnecessary pain and suffering. They're more damaging than copyright.

You actually don’t know that. The question would be, what proportion of human experiments are successful, and you don’t know the answer to that question, so the victims of experiments could dwarf the beneficiaries of successful research. That’s always the hard thing with basic utilitarianism.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#519

Earlier quoted context omitted.

> LibGen gives you access to a much smaller body of works than either of those. > Just go to a real library. The thrill of waiting a week for a book to arrive or navigating the labyrinthine interlibrary loan system is truly a privilege that many can afford. And who needs instant access to knowledge when you can have the pleasure of paying for shipping or commuting to a physical library? It's also fascinating that you…

Your library almost definitely offers digital loans as well.

Seeing the high prices they are charged for a digital licence which expires after a fairly small number of loans, I feel it'd be better for my library if I pirate when possible. Save those limited loans for someone who prefers/needs them.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#520

Earlier quoted context omitted.

" If you plug a laptop into a closet at MIT to download some scientific papers you forfeit your life." This is exactly what I immediately thought while reading the article. It almost feels like the legal system only punishes general public, while most of these guys are above it.

> MIT I think Aaron Swartz went to Harvard, not MIT https://en.wikipedia.org/wiki/United_States_v._Swartz

Yes, he went to Harvard; the laptop was plugged in at MIT using his Harvard Fellow credentials to access JSTOR.
Post reply on HN