Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

141–150 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#141

Earlier quoted context omitted.

True but lets take examples one by one to see what we can learn : Spotify was doing illegal things until they made a deal to become legal and not to be trialed over what they done. Seems like business deals is what saved them, not regulatory capture (the regulations around IP for music pre existed Spotify)

Sure that is what saved them initially, but following that early 2010’s period of hemorrhaging money and eventual recovery, then they started digging that moat. https://www.politico.com/story/2015/04/spotify-washington-lo... https://www.opensecrets.org/federal-lobbying/clients/issues?...

Very interesting thank you for the links. I'm not knowledgeable in the Music Modernization Act, but maybe some of this lobbying is to avoid being sued rather than building legal long term moat

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#142
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

" If you plug a laptop into a closet at MIT to download some scientific papers you forfeit your life." This is exactly what I immediately thought while reading the article. It almost feels like the legal system only punishes general public, while most of these guys are above it.

Wilhoit’s law:

> There must be in-groups whom the law protects but does not bind, alongside out-groups whom the law binds but does not protect.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#143
post #58

We all like hating big corporations, especially Meta, and people seem to use this as an opportunity to advocate for punishing them. I think it's wiser to advocate for changing our IP laws.

They broke the law and should be punished for that. Whether the law should change is a separate discussion.

Also, change the law so this is legal for poor meta? smh..

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#144
post #96
post #58

We all like hating big corporations, especially Meta, and people seem to use this as an opportunity to advocate for punishing them. I think it's wiser to advocate for changing our IP laws.

I truly hope Meta has a serious security issue that burns their company to the ground. That said, I want them to burn for the right reasons. Downloading data that should be available to the public is not one of them.

Exactly. Everyone should have the right to have access to this.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#145
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

I guess the solution is to create a shell company for your illegal activities?

You must be new to billionaire business practices: break the rules first, ask for forgiveness later.

By the time the cheque comes, your illicit venture either went bust or you built a bilion dollar empire capable of buying the best lawyers and lobbying to walk away clean.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#146

Earlier quoted context omitted.

Cain was severely punished. וְעַתָּ֖ה אָר֣וּר אָ֑תָּה מִן־הָֽאֲדָמָה֙ אֲשֶׁ֣ר פָּצְתָ֣ה אֶת־פִּ֔יהָ לָקַ֛חַת אֶת־דְּמֵ֥י אָחִ֖יךָ מִיָּדֶֽךָ׃ Therefore, you shall be more cursed than the ground, which opened its mouth to receive your brother’s blood from your hand. https://www.sefaria.org/Genesis.4.12

Jack the Ripper killed people and got away with it!!! Happy?

Better! Thanks.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#147
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

> Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. I don't understand why you wouldn't just buy copies of the books. Seems like such a relatively inexpensive way to strengthen your legal case.

Pretty sure that even if you gave a purchasing team enough money for retail price and a list of all books ever published, they wouldn't be able to buy even a quarter of them.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#148

Earlier quoted context omitted.

For creating a backup of library genesis. No. They should be awarded a philanthropic prize.

There's evidence of them seeding back as little as possible. I'm not sure how that's "creating a backup".

They're talking about creating and releasing Llama...not seeding the torrent

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#149
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

Something to understand about capitalist competition (also in politics) is that it's a war. Not one with guns and bombs, but more like a cold war, with espionage and hacking and just generally doing anything you can to gain an advantage without bringing negative consequences on yourself.

The limit is what you can actually get away with, not what the rules say you can get away with, and the system aggressively selects players who recognize this. It's amoral - there is no "ought", only "is". An actor gets punished or not, with absolutely no regard to whether it "should" get punished. One thing is consistent: following the rules as written means you lose.

You can see it in Y Combinator (and other) startups. The biggest ex-startups are things like AirBNB (hotels but we don't follow the rules but we don't get punished for not following them) and Uber (taxis but we don't follow the rules but we don't get punished for not following them).

One way to not get punished for not following the rules is to invent a variation of the game where the rules haven't been written yet. I again refer you to AirBNB and Uber; Omegle also comes to mind, although they didn't monetize.

Viewed in this light, Aaron Swartz's mistake was not the part where he downloaded journal articles, but the part where he got caught downloading journal articles. Shadow library sites are doing the same thing, minus the getting caught. So are Meta and Google and OpenAI. sci-hub is only involved in a lawsuit because it got caught and is now in the stage where it finds out whether it gets punished or not.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#150
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

> Google itself got big by indexing other people's data without compensation Wrong. a) Robots.txt which defines what content you wish to make available to third parties predates every search engine including Google. Web site owners chose to make it available to Google and search engines have respected their wishes despite it not being in their best interest. b) The difference here is that OpenAI, Meta etc have not ev…

Robots.txt is irrelevant after hiQ Labs v. LinkedIn (2019)
Post reply on HN