Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

131–140 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#131
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

> Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. I don't understand why you wouldn't just buy copies of the books. Seems like such a relatively inexpensive way to strengthen your legal case.

thanks to the byzantine copyright system, you can't easily do it. Plus, just speculating, but maybe by paying, it establishes "consideration" for some implied contract? "You implicitly entered a contract with us by purchasing the book, then violated the contract by 'distributing' the material for commercial use" ?

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#132
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

> Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. I don't understand why you wouldn't just buy copies of the books. Seems like such a relatively inexpensive way to strengthen your legal case.

Buying the books won't automatically give you permission to use the content commercially

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#133
post #51

It really makes you think about those crazy internet folks from back in the day who thought copyright law was too strict and that restricting humanity to knowledge in such a way was holding us all back for the benefit of a tiny few.

The more concerning thing is that the best thing these overpaid people could come up with was.. download the torrent, like everyone else. Here you are, billions of resources, and no one is willing to spend a part of it to at least digitize some new data? Like even Google did?

I think they are morally required to improve the current state.

- Seed the torrent and publicly promote piracy pushing lawmakers.

- Contribute with digitisation and open access like Google did in the past.

- Make the part of their dataset that was pirated publicly accessible.

- Fight stupid copyright laws. I can't believe that copyright lasts more than 20 years. No field moves that slowly, and there should be tighter limits on faster moving fields.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#134
post #110

Earlier quoted context omitted.

> Web site owners chose to make it available to Google. Strong disagree. Since robots.txt is optional and the default is "crawl me as you please", website owners don't "choose to make it available", they just don't choose to make it non-available.

That's a functionally meaningless distinction. If you setup a web server that responds to requests, then you're choosing to make content available because your server can choose to not respond to requests. The entire protocol includes mechanisms to negotiate access.

Granting access and granting right to redistribute (even just title + snippet) and use your content commercially are two completely different things.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#135

Earlier quoted context omitted.

People criming in the past is not an excuse for companies committing crimes today. You’re excusing lawlessness. Cain killed Abel and got away with it!! I can kill someone today too!!!

Cain was severely punished. וְעַתָּ֖ה אָר֣וּר אָ֑תָּה מִן־הָֽאֲדָמָה֙ אֲשֶׁ֣ר פָּצְתָ֣ה אֶת־פִּ֔יהָ לָקַ֛חַת אֶת־דְּמֵ֥י אָחִ֖יךָ מִיָּדֶֽךָ׃ Therefore, you shall be more cursed than the ground, which opened its mouth to receive your brother’s blood from your hand. https://www.sefaria.org/Genesis.4.12

Jack the Ripper killed people and got away with it!!! Happy?

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#136
post #8

Really curious what the judges are going to do here. Horse has functionally bolted on this already I’m guessing slap on wrist despite courts going after individual for a couple of movies torrented pretty hard

Is there any other possible outcome than a fine? That too one which will not really affect Meta's overall earnings

Ideally we have a conversation about how we as society have ended up in a situations where we have a two tier justice system.

At a minimum the starting point of discussion here should be that if life ruining $80,000 per item is an acceptable fine for individuals then why is it not the same for corporations. Which would probably get you a number in the trillions at which point we could have a discussion about reforming this entire system.

But yes realistically slap on wrist is what is going to happen here.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#137
post #110

Earlier quoted context omitted.

> Web site owners chose to make it available to Google. Strong disagree. Since robots.txt is optional and the default is "crawl me as you please", website owners don't "choose to make it available", they just don't choose to make it non-available.

That's a functionally meaningless distinction. If you setup a web server that responds to requests, then you're choosing to make content available because your server can choose to not respond to requests. The entire protocol includes mechanisms to negotiate access.

This is a meaningless simplification. In this framework "robots.txt" has no role, because your server "can choose" not to respond. Heck, even DDOS is fine, because "protocol"

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#138
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

I think if Google attempted to download the entirety of JSTOR with the express intent of making the full dataset freely available, then Google would also face legal consequences. It's true, and relevant, that Google would feel those consequences much less sharply than Swartz did.

Don't buy into the rhetoric and call it "consequences". It's always a choice to sue, a choice to prosecute, and this would be true even if these choices were made consistently and impartially (which they certainly aren't).

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#139
post #50

Good, we know it. Nothing will happen, because nothing happens to billionaires and their companies. Musk is proving it every day now.

This is why we need to abolish the government. If the government doesn't have any power, they can't do preferential treatment to their cronies.

Enough with laws for thee but not for me!

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#140
post #66
post #56

Earlier quoted context omitted.

> Google itself got big by indexing other people's data without compensation Weird framing given how much value was and is still placed on Google driving traffic to you

For Google's case the order was reversed. Google used to send customers to your site. Now they try to show you the information on their site so that the customer doesn't need to go to your site.

Unless your site is an SEO cesspool serving THEIR ads, then you do end up there.
Post reply on HN