Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

71–80 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#71

If you're an author with a book likely to have be hoovered up, I wonder what you'd get from the fb models if you asked "complete this in the style of [author] in [book]: [quite a long excerpt]" If you get a direct quote then you're good with your claim, surely.

I believe that is part of this lawsuit pretty much

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#72
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

" If you plug a laptop into a closet at MIT to download some scientific papers you forfeit your life."

This is exactly what I immediately thought while reading the article. It almost feels like the legal system only punishes general public, while most of these guys are above it.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#73
post #41

Considering prices for single work, this must be multi-billion dollar compensation. Take for example 675k paid for 31 songs. So 20k a song. If we estimate book to be say 10MB that would 8 million works. So I think reasonable compensation is something along 163 billion. Not even 10 years of net income. Which I think is entirely fair punishment.

For creating a backup of library genesis. No. They should be awarded a philanthropic prize.

There's evidence of them seeding back as little as possible. I'm not sure how that's "creating a backup".

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#74
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

Comprehensive intellectual property needs to happen for the modern (digital) era. Basically the entire legal system needs to be retooled and rethought for computers.

Looks like the entire legal system is being retooled at the moment.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#75
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

the english empire once tried to mantain a monopoly over steam loom machines the americans cheated their way to competition, heck, even before that, the english empire got jumpstarted by stealing gold from the spanish (who were themselves exploiting it away from aztec and other mexican natives) I'm saying it's business as usual, but also, culture doesn't work like tangible physical widgets so we must stop letting a f…

Hollywood became popular for filmmaking because they were literally the opposite side of the country from Thomas Edison and his patents...

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#76
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

> Google itself got big by indexing other people's data without compensation

Wrong.

a) Robots.txt which defines what content you wish to make available to third parties predates every search engine including Google. Web site owners chose to make it available to Google and search engines have respected their wishes despite it not being in their best interest.

b) The difference here is that OpenAI, Meta etc have not even tried to honour the wishes of copyright holders. They just considered everything as theirs.

c) Google grew big because it had no ads, fast interface and PageRank was significantly better. It wasn't because it had the most comprehensive index.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#78
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

> Google itself got big by indexing other people's data without compensation Wrong. a) Robots.txt which defines what content you wish to make available to third parties predates every search engine including Google. Web site owners chose to make it available to Google and search engines have respected their wishes despite it not being in their best interest. b) The difference here is that OpenAI, Meta etc have not ev…

To your first point, the op said without compensation, not without permission.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#79
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

> Google itself got big by indexing other people's data without compensation Wrong. a) Robots.txt which defines what content you wish to make available to third parties predates every search engine including Google. Web site owners chose to make it available to Google and search engines have respected their wishes despite it not being in their best interest. b) The difference here is that OpenAI, Meta etc have not ev…

point c is wrong...they had ads since the original yahoo contract....

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#80
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

> Google itself got big by indexing other people's data without compensation Wrong. a) Robots.txt which defines what content you wish to make available to third parties predates every search engine including Google. Web site owners chose to make it available to Google and search engines have respected their wishes despite it not being in their best interest. b) The difference here is that OpenAI, Meta etc have not ev…

> Web site owners chose to make it available to Google.

Strong disagree. Since robots.txt is optional and the default is "crawl me as you please", website owners don't "choose to make it available", they just don't choose to make it non-available.

Post reply on HN