Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

201–210 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#201
post #166

Earlier quoted context omitted.

Airbnb and Uber have showed us that laws matter only to the extent that the political will to enforce them exists. Throw enough lawyers and lobbying money at the problem and the laws can simply be re-written to be friendlier to your business model.

The reason there was no political will to punish Airbnb and Uber for violating the law was that initially they were subsidized with VC money and so were able to undercut traditional hotels and taxis on price. In the world of tradable goods, pricing below cost with the intent of putting competition out of business so you can raise prices later is known as "dumping" and is itself illegal.

The reason is that everyone who was supposed to do something about it was "subsidized with VC money".

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#202
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

The most outrageous thing about the whole story is that smart people (like here and not only) knew this all since day one. They been uncovering this the whole time.

And in their face, with all the fierce ignorance, broligarchs deny, evade and totally pretend this never happened. The most non open company of all even went to lengths to accuse others of stealing their IP - not theirs to begin with.

Just think of it - why did all major content platforms closed their APIs the day after GPT-2 got the word going…? Cause they knew all this very well - the content is precious and needed. They been doing it all along. Distilling the essence of world’s writing and digital imagery they had no right to.

We have a saying where I come from - no mercy for the chicken, no laws for the millions. I thought it was a local thing at first, it turned is how the world goes. Nothing new under the sun, indeed.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#203
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

Yes. And the problem here isn't that companies get away with doing things like this, the problem is that individuals don't. Attempting to lock information behind a nightmarish legal system is the problem. I'm pretty much at the point now where I don't buy the "copyright incentivizes creation" argument any more. Copyright, like advertising, incentivizes creation by enormous corporations, but also like advertising it i…

Also, in Canada, it's basically impossible to protect your IP as an individual due to the astronomical cost and lack of options to recover that cost. So copyright will never incentivize my creations, or those of any small creator.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#205
post #123

Libgen is a civilizational project that should be endorsed, not prosecuted. I hope one day people will look at it and think how stupid we were today to shun the largest collection of literary works in human history.

I think you’re overstating its importance. The internet already makes it possible to order almost any book in existence and have it arrive at your doorstep within a week or so, or often on your ebook reader instantly. And your local library probably participates in an interlibrary loan system that lets you request any book held by any library in the country for free . LibGen gives you access to a much smaller body of…

[flagged]

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#206
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

the english empire once tried to mantain a monopoly over steam loom machines the americans cheated their way to competition, heck, even before that, the english empire got jumpstarted by stealing gold from the spanish (who were themselves exploiting it away from aztec and other mexican natives) I'm saying it's business as usual, but also, culture doesn't work like tangible physical widgets so we must stop letting a f…

I don't think I've heard the term "English empire". Is it an attempt by the Scottish to pretend they weren't involved?

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#207
post #157
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

The thing is Google, meta and YouTube weren't giant entities when they did this stuff. I think it's good no one cracked down on them for copyright stuff. Now they're developing an LLM that will generate potentially trillions in value to humanity and looks like they're not exactly playing by the rules. But I prefer looser intellectual property rights anyway so Im ok with it

>But I prefer looser intellectual property rights anyway so Im ok with it

I think more people, potentially anyways, would feel similar to to this if it applied even somewhat equally.

Instead, companies can seemingly do whatever they please whereas lawyers will send letters to your home for downloading a single episode of game of thrones.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#208

Earlier quoted context omitted.

I think if Google attempted to download the entirety of JSTOR with the express intent of making the full dataset freely available, then Google would also face legal consequences. It's true, and relevant, that Google would feel those consequences much less sharply than Swartz did.

Don't buy into the rhetoric and call it "consequences". It's always a choice to sue, a choice to prosecute, and this would be true even if these choices were made consistently and impartially (which they certainly aren't).

I wasn't meaning to attach a pejorative to "consequences", but the word does typically have that meaning so you're right to call me out. Perhaps "resulting legal issues" would be a better way to put it.

For the record, I think the consequence was grossly disproportionate to the action.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#209

Libgen is a civilizational project that should be endorsed, not prosecuted. I hope one day people will look at it and think how stupid we were today to shun the largest collection of literary works in human history.

Libgen turns into a problem when you have a company developing generative AI with it, either giving money to GPU manufacturers or themselves with paid services (see OpenAI)

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#210
post #123

Libgen is a civilizational project that should be endorsed, not prosecuted. I hope one day people will look at it and think how stupid we were today to shun the largest collection of literary works in human history.

I think you’re overstating its importance. The internet already makes it possible to order almost any book in existence and have it arrive at your doorstep within a week or so, or often on your ebook reader instantly. And your local library probably participates in an interlibrary loan system that lets you request any book held by any library in the country for free . LibGen gives you access to a much smaller body of…

There are a whole lot of books that are out of print, and if a book went out of print before ebooks were a thing, it probably doesn't have a legal digital edition either.
Post reply on HN