Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

351–360 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#351

Earlier quoted context omitted.

They should sell for a price it would make pirating it pointless. Like what Spotify or Netflix did to audio and visual content. Then they can use the exposure to find other ways to make money.

Or if you don't agree with the price, do not buy it. You are NOT entitled to entertain yourself in any way you want. (unless it's funded by taxes etc, in which case... okay, it's open to discussion.) Look, let's be honest - what gives you or others the right to steal from others?

That's what I do, personally.

But you call it "stealing," others call it "copying."

Stealing takes, from someone, something they own.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#352

I strongly urge people to read Thomas Babington Macaulay's speeches on copyright, its aims, terms, and hazards. Very well reasoned and explained. In particular, people often cited the case of authors who had died leaving a family in destitution, and claimed that copyright extension would be a fair way of preventing this, but in most cases the remaining family had never held the copyright; the author had initally sold…

This one example does not make stealing acceptable which is what you’re implying.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#353

Earlier quoted context omitted.

True but lets take examples one by one to see what we can learn : Spotify was doing illegal things until they made a deal to become legal and not to be trialed over what they done. Seems like business deals is what saved them, not regulatory capture (the regulations around IP for music pre existed Spotify)

Same would be the case for YouTube. Google case was different in that AFAIK there wasn't any obvious legal problem with indexing, and, back then, they were actually doing everyone a favor. Hardly anyone had any issue with Google search until the time when news media screwed themselves over by going all in on ads, overdoing it, then trying to bring back the paywall, only to realize no one is actually browsing their si…

And then news media in Canada got even worse a few years ago. They demanded the government make a law so that when Google or Facebook even links to an article, they must pay the news org. Google decided to pay the link tax. Facebook decided to block all news links. From talking to people, most think that Facebook is the villain, in reality it's the news orgs in collusion with state power. https://en.wikipedia.org/wiki/Online_News_Act , https://www.justice.gc.ca/eng/csj-sjc/pl/charter-charte/c18_...

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#354
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

" If you plug a laptop into a closet at MIT to download some scientific papers you forfeit your life." This is exactly what I immediately thought while reading the article. It almost feels like the legal system only punishes general public, while most of these guys are above it.

> the legal system only punishes general public, while most of these guys are above it

It’s because the legal system is not about justice, it’s about money

Most people can’t afford lawyers or expensive legal battles

On the other hand, individuals and organizations with a lot of money get to weaponize and exploit the legal system to their advantage

“To my friends, anything; to my enemies, the law”

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#355
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

" If you plug a laptop into a closet at MIT to download some scientific papers you forfeit your life." This is exactly what I immediately thought while reading the article. It almost feels like the legal system only punishes general public, while most of these guys are above it.

if you get a group of people and call it an llc then criminal elements are largely eliminated.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#356

Earlier quoted context omitted.

They should sell for a price it would make pirating it pointless. Like what Spotify or Netflix did to audio and visual content. Then they can use the exposure to find other ways to make money.

Or if you don't agree with the price, do not buy it. You are NOT entitled to entertain yourself in any way you want. (unless it's funded by taxes etc, in which case... okay, it's open to discussion.) Look, let's be honest - what gives you or others the right to steal from others?

> what gives you or others the right to steal from others?

I think technically it's copying more than stealing

Like if you could wait for someone to design and build a car and then CTRL+C/V it for yourself (is it possible to steal in a post-scarcity society?)

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#357

I don't understand why it's even a question that Meta trained their LLM on copyrighted material. They say so in their paper! Quoting from their LLaMMa paper [Touvron et al., 2023]: > We include two book corpora in our training dataset: the Gutenberg Project, [...], and the Books3 section of ThePile (Gao et al., 2020), a publicly available dataset for training large language models. Following that reference: > Books3…

There are two different things when it comes to discussing training LLM's on "copyright" protected data, and I almost never see people differentiate. 1.) Training on copyright that is publicly available. You write a poem and publish it online for the world to read. That is your IP, no one else can take it an sell it, but they are free to read and be inspired by it. The legalitly of training on this is in the courts,…

good distinction

IMO there's a hack about this,

authors can claim that they allow for public use unless it's used for training LLMs. And all of training work would fall under 2 because they would be used against the copyright.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#358

It really makes you think about those crazy internet folks from back in the day who thought copyright law was too strict and that restricting humanity to knowledge in such a way was holding us all back for the benefit of a tiny few.

Probably the single biggest thing I learned growing up is that you can safely live by "Everyone is in it for themselves".

It's incredibly rare to find people who hold ideals that are detrimental to their own life.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#359
post #123

Earlier quoted context omitted.

I think you’re overstating its importance. The internet already makes it possible to order almost any book in existence and have it arrive at your doorstep within a week or so, or often on your ebook reader instantly. And your local library probably participates in an interlibrary loan system that lets you request any book held by any library in the country for free . LibGen gives you access to a much smaller body of…

> LibGen gives you access to a much smaller body of works than either of those. > Just go to a real library. The thrill of waiting a week for a book to arrive or navigating the labyrinthine interlibrary loan system is truly a privilege that many can afford. And who needs instant access to knowledge when you can have the pleasure of paying for shipping or commuting to a physical library? It's also fascinating that you…

Yes, publishers don’t pay authors as much as they deserve, but LibGen pays them literally nothing. Authors tend to love libraries but hate piracy. Why? Because earning something is better than earning nothing.

Have you ever submitted an ILL request? It’s extremely simple. Many library systems even integrate with WorldCat, so submitting a request for any book just takes a few clicks.

I’m mostly speaking about people in the US. Every single county in the entire country has a public library. Almost all of them have ILL.

I think equity is a fair argument for the existence of services like LibGen in many parts of the world, but the reality is that almost everyone using a book piracy sites in a first-world country is using it to pirate an in-print book that they just don’t want to go to the trouble of borrowing or buying.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#360

If you're an author with a book likely to have be hoovered up, I wonder what you'd get from the fb models if you asked "complete this in the style of [author] in [book]: [quite a long excerpt]" If you get a direct quote then you're good with your claim, surely.

That's the NYT's case. Not necessarily very strong. https://www.techdirt.com/2024/03/05/openais-motion-to-dismis...
Post reply on HN