Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

171–180 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#171
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

> Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. I don't understand why you wouldn't just buy copies of the books. Seems like such a relatively inexpensive way to strengthen your legal case.

Buying a copy of the book doesn’t grant you the right to copy it. That is what copyright is for.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#172
post #58

We all like hating big corporations, especially Meta, and people seem to use this as an opportunity to advocate for punishing them. I think it's wiser to advocate for changing our IP laws.

I think most of the public is probably in favor of stronger IP laws now that big corps are threatening to make them jobless with IP-disrespecting AIs

Something tells me stronger IP laws will be drafted by holders of that IP, with little if any regard to the potential for job losses for regular people from AI.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#173
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

> If you plug a laptop into a closet at MIT to download some scientific papers you forfeit your life.

Just to point, but the material in question was public domain, so nobody had even a copyrights claim over it.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#174
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

" If you plug a laptop into a closet at MIT to download some scientific papers you forfeit your life." This is exactly what I immediately thought while reading the article. It almost feels like the legal system only punishes general public, while most of these guys are above it.

>This is exactly what I immediately thought while reading the article. It almost feels like the legal system only punishes general public, while most of these guys are above it.

Welcome to the modern day aristocracy. Not only what you mentioned, this world is also divided into a group of insider who can get capital from 0 - 2%, while rest of us has a cost of 17%, 22% or 30%?

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#175
post #41

Considering prices for single work, this must be multi-billion dollar compensation. Take for example 675k paid for 31 songs. So 20k a song. If we estimate book to be say 10MB that would 8 million works. So I think reasonable compensation is something along 163 billion. Not even 10 years of net income. Which I think is entirely fair punishment.

Meta argues that it's fair use, and that they just downloaded, and never seeded, all the torrents.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#176

Earlier quoted context omitted.

For creating a backup of library genesis. No. They should be awarded a philanthropic prize.

There's evidence of them seeding back as little as possible. I'm not sure how that's "creating a backup".

In that case they should also be sued for not complying with bittorrent's tit-for-tat ethiquette. Leechers should be punished. :)

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#177
post #50

Good, we know it. Nothing will happen, because nothing happens to billionaires and their companies. Musk is proving it every day now.

This is why we need to abolish the government. If the government doesn't have any power, they can't do preferential treatment to their cronies. Enough with laws for thee but not for me!

The problem is precisely that those billionaires are too powerful. If anything, we need to abolish the billionaires.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#178
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

I think if Google attempted to download the entirety of JSTOR with the express intent of making the full dataset freely available, then Google would also face legal consequences. It's true, and relevant, that Google would feel those consequences much less sharply than Swartz did.

Google book search was declared fair use and copyright holders ended up having to explicitly request removal of their works.

Apparently he would have gotten away with downloading the JSTOR database if he made it clear that he intended to only publish half of each paper.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#179
post #60
post #58

We all like hating big corporations, especially Meta, and people seem to use this as an opportunity to advocate for punishing them. I think it's wiser to advocate for changing our IP laws.

First punish them. Then change the laws.

I bet you and my "first build the product, then worry about security" manager would get along.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#180
post #125

Earlier quoted context omitted.

I wasn't aware. Can you please update Wikipedia then: https://en.wikipedia.org/wiki/Robots.txt Maybe also get Google to update their docs: https://developers.google.com/search/docs/crawling-indexing/...

their own docs also specify that the robots.txt does not stop indexing or showing up in search, they even bolded it "it is not a mechanism for keeping a web page out of Google" https://developers.google.com/search/docs/crawling-indexing/...

The only way for links to appear in a Google search would be to host a public resource, that is linked from another public resource.

If you have specified in your robots.txt that you do not want the page(s) or directories ingested then only the url is indexed (if it is linked from another page). It does prevent the public display of the content of a page and creation description/summary.

https://support.google.com/webmasters/answer/7489871?hl=en

Post reply on HN