Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

101–110 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#101
post #33
post #23

[flagged]

Not that I have any particular sympathy for the guy, but could we keep the tone a bit more civil around here? HN is one of the few bastions of grounded discussion on the internet, and I’d prefer to keep it that way.

> HN is one of the few bastions of grounded discussion on the internet

Hacker News has its fair share of irrationality.

I submitted an article about efforts to undermine Wikipedia: https://news.ycombinator.com/item?id=42962971

But it was flagged and locked to commenting.

Wikipedia is one of the internet's greatest projects. Hacker News apparently doesn't have the stomach to discuss the threats to Wikipedia.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#102
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

I guess the solution is to create a shell company for your illegal activities?

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#104

Earlier quoted context omitted.

Wrong. Google ignores robots.txt entirely

I wasn't aware. Can you please update Wikipedia then: https://en.wikipedia.org/wiki/Robots.txt Maybe also get Google to update their docs: https://developers.google.com/search/docs/crawling-indexing/...

It must be nice to believe everything people say by default... ;)

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#105
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

the english empire once tried to mantain a monopoly over steam loom machines the americans cheated their way to competition, heck, even before that, the english empire got jumpstarted by stealing gold from the spanish (who were themselves exploiting it away from aztec and other mexican natives) I'm saying it's business as usual, but also, culture doesn't work like tangible physical widgets so we must stop letting a f…

Why do I get sued when I share some BitTorrents but $bigcorp can just do it with 1000 scale without problems?

The issue here is not copyright/patents/etc - the issue is that the law is applied selectively — the issue is that Aaron Schwartz is dead for sharing knowledge with the public and Zuccborg is a billionaire building his torment nexus

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#106
post #89

Earlier quoted context omitted.

" If you plug a laptop into a closet at MIT to download some scientific papers you forfeit your life." This is exactly what I immediately thought while reading the article. It almost feels like the legal system only punishes general public, while most of these guys are above it.

The legal system is built to favor large corps and capital owners. See Katharina Pistor books for instance.

I think it’s the other way around. Those large entities break all the same laws and rules as others and then get to the point where they can influence the creation of a regulatory moat around themselves to prevent competitors from taking the same path as them.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#107
post #91

Earlier quoted context omitted.

Which was arguably more innocent — scientific papers.

Meta is not “innocent”, and comparing this instance with Swartz is a huge offense to his legacy.

I think comparing it is reasonable and valid. Equaling it would be incorrect. What Meta is (allegedly, likely) doing here is several orders of magnitude worse, in scale and intention. I'd say both ethical and probably juristical.

But just because the scale and intention are different, does not mean we cannot compare both cases. They are not equal, far from it. But they are compareable.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#108
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

> Google itself got big by indexing other people's data without compensation Wrong. a) Robots.txt which defines what content you wish to make available to third parties predates every search engine including Google. Web site owners chose to make it available to Google and search engines have respected their wishes despite it not being in their best interest. b) The difference here is that OpenAI, Meta etc have not ev…

a) If you don't have a robots.txt, you're indexed by default. It's opt-out, not opt-in. If you do nothing, you're being indexed.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#109
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

the english empire once tried to mantain a monopoly over steam loom machines the americans cheated their way to competition, heck, even before that, the english empire got jumpstarted by stealing gold from the spanish (who were themselves exploiting it away from aztec and other mexican natives) I'm saying it's business as usual, but also, culture doesn't work like tangible physical widgets so we must stop letting a f…

Interesting, if we're to trust what NotOpenAI and Facebook say about their IP, the US should pay the UK reparations for IP theft based on textile industry profits starting in the 1850s until today?

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#110

Earlier quoted context omitted.

> Google itself got big by indexing other people's data without compensation Wrong. a) Robots.txt which defines what content you wish to make available to third parties predates every search engine including Google. Web site owners chose to make it available to Google and search engines have respected their wishes despite it not being in their best interest. b) The difference here is that OpenAI, Meta etc have not ev…

> Web site owners chose to make it available to Google. Strong disagree. Since robots.txt is optional and the default is "crawl me as you please", website owners don't "choose to make it available", they just don't choose to make it non-available.

That's a functionally meaningless distinction. If you setup a web server that responds to requests, then you're choosing to make content available because your server can choose to not respond to requests. The entire protocol includes mechanisms to negotiate access.
Post reply on HN