Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…
> Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. I don't understand why you wouldn't just buy copies of the books. Seems like such a relatively inexpensive way to strengthen your legal case.
Meta torrented & seeded 81.7 TB dataset containing copyrighted data
171–180 of 981 posts
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#172We all like hating big corporations, especially Meta, and people seem to use this as an opportunity to advocate for punishing them. I think it's wiser to advocate for changing our IP laws.
I think most of the public is probably in favor of stronger IP laws now that big corps are threatening to make them jobless with IP-disrespecting AIs
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#173Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…
Just to point, but the material in question was public domain, so nobody had even a copyrights claim over it.
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#174Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…
" If you plug a laptop into a closet at MIT to download some scientific papers you forfeit your life." This is exactly what I immediately thought while reading the article. It almost feels like the legal system only punishes general public, while most of these guys are above it.
Welcome to the modern day aristocracy. Not only what you mentioned, this world is also divided into a group of insider who can get capital from 0 - 2%, while rest of us has a cost of 17%, 22% or 30%?
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#175Considering prices for single work, this must be multi-billion dollar compensation. Take for example 675k paid for 31 songs. So 20k a song. If we estimate book to be say 10MB that would 8 million works. So I think reasonable compensation is something along 163 billion. Not even 10 years of net income. Which I think is entirely fair punishment.
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#176Earlier quoted context omitted.
For creating a backup of library genesis. No. They should be awarded a philanthropic prize.
There's evidence of them seeding back as little as possible. I'm not sure how that's "creating a backup".
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#177Good, we know it. Nothing will happen, because nothing happens to billionaires and their companies. Musk is proving it every day now.
This is why we need to abolish the government. If the government doesn't have any power, they can't do preferential treatment to their cronies. Enough with laws for thee but not for me!
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#178Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…
I think if Google attempted to download the entirety of JSTOR with the express intent of making the full dataset freely available, then Google would also face legal consequences. It's true, and relevant, that Google would feel those consequences much less sharply than Swartz did.
Apparently he would have gotten away with downloading the JSTOR database if he made it clear that he intended to only publish half of each paper.
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#179We all like hating big corporations, especially Meta, and people seem to use this as an opportunity to advocate for punishing them. I think it's wiser to advocate for changing our IP laws.
First punish them. Then change the laws.
Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data
#180Earlier quoted context omitted.
I wasn't aware. Can you please update Wikipedia then: https://en.wikipedia.org/wiki/Robots.txt Maybe also get Google to update their docs: https://developers.google.com/search/docs/crawling-indexing/...
their own docs also specify that the robots.txt does not stop indexing or showing up in search, they even bolded it "it is not a mechanism for keeping a web page out of Google" https://developers.google.com/search/docs/crawling-indexing/...
If you have specified in your robots.txt that you do not want the page(s) or directories ingested then only the url is indexed (if it is linked from another page). It does prevent the public display of the content of a page and creation description/summary.