Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

261–270 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#262

Best way to "punish" Meta is to slash the Gordian knot and abolish copyright. Level the playing field, incrementally, for everyone else who isn't a trillion-dollar corporation. The alternative is a futile legalistic attack against a monopoly entity too powerful to be meaningfully punished. That won't accomplish anything useful. It would, rather, help cement this status quo, where copyright infringement is selectively…

What's the alternative to copyright then? Anything I create will be instantly reproduced and sold for less than I can afford to by some entity far larger and more efficient than me.

> Level the playing field, incrementally, for everyone else who isn't a trillion-dollar corporation.

There is no level playing field when you have individuals and trillion-dollar companies in the same market.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#263
post #241
post #209

Earlier quoted context omitted.

Libgen turns into a problem when you have a company developing generative AI with it, either giving money to GPU manufacturers or themselves with paid services (see OpenAI)

…why? Will people buy less books because we have intuitive algorithms trained on old books? Personally, I strongly believe that the aesthetic skills of humanity are one of our most advanced faculties — we are nowhere close to replacing them with fully-automated output, AGI or no.

old books? i can imagine the shit/hallucinated-like generative AI we would have if the training weight was restricted to public domain stuff...

i think when chatGPT was around version 2 or 3, i had extracted almost 2 pages (without any alteration from the original) with questions that considered the author from this book here, https://www.amazon.com/Loneliness-Human-Nature-Social-Connec...

now it's up to you to think this is okay... but i bet you are no author

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#264
post #246

How about a consequentialist argument? In some fields, AI has already surpassed physicians in diagnosing illnesses. If breaking copyright laws allows AI to access and learn from a broader range of data, it could lead to earlier and more accurate diagnoses, saving lives. In this case, the ethical imperative to preserve human life outweighs the rigid enforcement of copyright laws.

There’s nothing particular to AI about your comment, it’s a general downside of IP.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#265

Earlier quoted context omitted.

I guess the solution is to create a shell company for your illegal activities?

The modern solution has been to grow so fast that by the time anyone can go after you legally you've already amassed so much money/power that you can have the laws rewritten (or at least enforced) around your existence. IMO part of the reason the SV tech bros are embracing right wing grift culture so publicly now is that this method, which had been serving them well for decades, doesn't really work without the infini…

That's why you should go straight to the treasury's RSS feeds.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#266

Earlier quoted context omitted.

There's evidence of them seeding back as little as possible. I'm not sure how that's "creating a backup".

They're talking about creating and releasing Llama...not seeding the torrent

A model is not a backup.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#268
post #157
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

The thing is Google, meta and YouTube weren't giant entities when they did this stuff. I think it's good no one cracked down on them for copyright stuff. Now they're developing an LLM that will generate potentially trillions in value to humanity and looks like they're not exactly playing by the rules. But I prefer looser intellectual property rights anyway so Im ok with it

Well, we'll see how will it generate value and for whom.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#269
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

" If you plug a laptop into a closet at MIT to download some scientific papers you forfeit your life." This is exactly what I immediately thought while reading the article. It almost feels like the legal system only punishes general public, while most of these guys are above it.

[deleted]

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#270
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

" If you plug a laptop into a closet at MIT to download some scientific papers you forfeit your life." This is exactly what I immediately thought while reading the article. It almost feels like the legal system only punishes general public, while most of these guys are above it.

It's an oligarchy, always has been. I don't know how colossal the pile of evidence supporting this has to get before people finally accept it.
Post reply on HN