Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

611–620 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#611

Earlier quoted context omitted.

I find this such a strange remark on this front. You got less than 1% of a book... from an author who has passed away... who wrote on a research topic that was funded by an institution that takes in hundreds of millions of dollars in federal grants each year... I'm not an author (although I do generate almost exclusively IP for a living) and I think this is about as weak a form of this argument as you possibly make.…

I think the key is to think through the incentives for future authors. As a thought experiment, say that the idea someday becomes mainstream that there is no reason to read any book or research publication because you can just ask an AI to describe and quote at length from the contents of anything you might want to read. In such a future, I think it's reasonable to predict that there would be less incentive to publis…

  there is no reason to read any book or research publication because you can just ask an AI to describe and quote at length from the contents of anything you might want to read
I think this is the fundamental misunderstanding at the heart of a lot of the anger over this, beyond the basic "corporations in general are out of control and living authors should earn a fair wage" points that existed before this.

You summarize well how we aren't there yet, but I'd say the answer to your final implied question is "not likely to change at all". Even when my fellow traitors-to-humanity are done with our cognitive AGI systems that employ intuitive algorithms in symphony with deliberative symbolic ones, at the end of the day, information theory holds for them just as much as it does for us. LLMs are not built to memorize knowledge, they're built to intuitively transform text -- the only way to get verbatim copies of "anything you might want to read" is fundamentally to store a copy of it. Full stop, end of story, will never not be true.

In that light, such a future seems as easy to avoid today as it was 5 years ago: not trivial, but well within the bounds of our legal and social systems. If someone makes a bot with copies of recent literature, and the authors wrote that lit under a social contract that promised them royalties, then the obvious move is to stop them.

Until then, as you say: only extremists and laymen who don't know better are using LLMs to replace published literature altogether. Everyone else knows that the UX isn't there, and the chance for confident error way too high.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#612
post #263

Earlier quoted context omitted.

old books? i can imagine the shit/hallucinated-like generative AI we would have if the training weight was restricted to public domain stuff... i think when chatGPT was around version 2 or 3, i had extracted almost 2 pages (without any alteration from the original) with questions that considered the author from this book here, https://www.amazon.com/Loneliness-Human-Nature-Social-Connec... now it's up to you to think…

I find this such a strange remark on this front. You got less than 1% of a book... from an author who has passed away... who wrote on a research topic that was funded by an institution that takes in hundreds of millions of dollars in federal grants each year... I'm not an author (although I do generate almost exclusively IP for a living) and I think this is about as weak a form of this argument as you possibly make.…

that was just a metaphor, you can ask your AI what's that or use way less energy and use Wikipedia's search engine... or do you think OpenAI first evaluates if the author is an independent developer &/or has died &/or was funded by a public university before adding the content to the training database? /s

and one thing is publishing a paper with jargon for academics, another is to simplify the results for the masses. there's a huge difference between finishing a paper and a book

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#613

The question is, if they could and would have paid for each book, would it be ok to train the LLM on them? I'm talking about prior books, I'm sure new books have language forbidding their use to train LLMs at the point of sale. But legally, how does using a book to train a LLM differ from a teacher learning from a book and teaching its contents to their pupils. Obviously, the LLM can do so at scale, but is there a le…

> The question is, if they could and would have paid for each book, would it be ok to train the LLM on them? Whether training on AI model on an array of diffentent works, many of which are copyright protected, is itself a copyright violation, in addition to or distinct from any copyright violation that goes on gathering the dataset for training (and separate from any copyright violation in the actual or intended use…

Thank you for a good answer.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#614
post #185
post #58

We all like hating big corporations, especially Meta, and people seem to use this as an opportunity to advocate for punishing them. I think it's wiser to advocate for changing our IP laws.

You're conflating different problems. Big corporations are too big, they should just not exist. When you have corporations more powerful than the government of the biggest states, it's a bug, not a feature. The IP laws may need rethinking. Saying that they should disappear because big corporations are above the law doesn't help, though. First kill the big corporations, then think about fair laws. Changing the law now…

> When you have corporations more powerful than the government of the biggest states, it's a bug, not a feature.

The only distinction between corporations and governments is one of them are morally bankrupt arbiters of force.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#615
post #598

Earlier quoted context omitted.

Several of those things aren't even necessarily illegal and are the sort of things they shouldn't have had have any reason to do unless they were being targeted by a media campaign or captured government. There is also some dispute about whether some of those even happened or are just mischaracterizations from the media campaign. It's like saying "well, they weren't only violating the taxi medallion cartel laws, they…

Move the goalposts any more and they’re going to be outside the stadium. What laws matter to you? I agree there are shit laws but why can uber break them with impunity but individuals are jailed for smoking some fun lettuce?

Because more money and special interests are behind fun lettuce smoking enforcement than local taxi companies could put behind protecting their own cartel from interlopers. If the taxi companies had more money to dump on politicians than is poured into drug enforcement, then the priorities would have changed.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#616
post #349
post #185

Earlier quoted context omitted.

You're conflating different problems. Big corporations are too big, they should just not exist. When you have corporations more powerful than the government of the biggest states, it's a bug, not a feature. The IP laws may need rethinking. Saying that they should disappear because big corporations are above the law doesn't help, though. First kill the big corporations, then think about fair laws. Changing the law now…

> First kill the big corporations, then think about fair laws. It's not possible to kill big corporations before fair laws, because as you said yourself "corporations are already above the law" Unfair laws don't apply to big corporations, they only apply to the people opposed to big corporations It's akin to hamstringing a horse and saying you'll fix it when they win

Anti-Trust laws are a little different though. It's specifically about bringing giant corps down a peg, and has been used multiple times against companies that otherwise skirt the law quite a bit.

Standard Oil, AT&T, the railroads, all thought they were above the law, for good reason, but they were all still broken.

Not going to happen for 4 years at least.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#617

Earlier quoted context omitted.

> If you reject Macaulay on copyright because he was an imperialist On the contrary I would argue that this is precisely why you SHOULD NOT take his opinion on copyright. One of the main outcomes of imperialism/colonization is denigrating/destroying/appropriating works of art, literature with the primary goal of subjugation, subversion and thereupon replacement of native culture/traditions/institutions. I did not quo…

> One of the main outcomes of imperialism/colonization is denigrating/destroying/appropriating works of art, literature with the primary goal of subjugation, subversion and thereupon replacement of native culture/traditions/institutions. Which is irrelevant to the question of whether copyright law within the country of England and within English culture is beneficial or not. It is the nature of racism that it bypasse…

Feels a little pat though doesn't it? If racism itself is necessarily defined by irrationality then you'd think the entire course of Western civilization would have gone a little differently. Not to mention, we have some pretty dark lessons from history already that are precisely the result of excessive rationality. One could easily demonstrate the "rationality" of a given colonial project, for example.

I'm not saying we need to choose between a broader humanism or rationality necessarily, but I just think it feels a little archaic Enlightenment-era thinking to reduce it down this particular way. Or just you know, its all Spock and no Kirk!

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#618

Earlier quoted context omitted.

Then it would need to be determined, whether that is the case or not. Did every single machine they used have the configuration for only leeching and no seeding? The company is liable for what its employees on the job. If only one employee was also seeding ... that could be a very interesting case.

> Did every single machine they used have the configuration for only leeching and no seeding? I would certainly assume so. It's incredibly obvious that's what you would want to do from a legal standpoint. > If only one employee was also seeding ... that could be a very interesting case. The torrenting wouldn't be done casually by employees acting on their own. And it's not like multiple employees are doing it simulta…

Did you not read the article? There are quotes from Meta employees doing exactly what you claim they wouldn't do.

> This is part of an official project. They'd spin up a machine just to download the torrent, being careful to disable seeding.

From the article:

> "Torrenting from a corporate laptop doesn’t feel right," Nikolay Bashlykov, a Meta research engineer, wrote in an April 2023 message, adding a smiley emoji. In the same message, he expressed "concern about using Meta IP addresses 'to load through torrents pirate content.'"

You also claim they would be "careful to disable seeding" but we know they did in fact seed (and anyone who uses private trackers knows they couldn't get away with leeching for very long before being kicked off):

> Meta also allegedly modified settings "so that the smallest amount of seeding possible could occur," a Meta executive in charge of project management, Michael Clark, said in a deposition.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#619

I don't understand why it's even a question that Meta trained their LLM on copyrighted material. They say so in their paper! Quoting from their LLaMMa paper [Touvron et al., 2023]: > We include two book corpora in our training dataset: the Gutenberg Project, [...], and the Books3 section of ThePile (Gao et al., 2020), a publicly available dataset for training large language models. Following that reference: > Books3…

There are two different things when it comes to discussing training LLM's on "copyright" protected data, and I almost never see people differentiate. 1.) Training on copyright that is publicly available. You write a poem and publish it online for the world to read. That is your IP, no one else can take it an sell it, but they are free to read and be inspired by it. The legalitly of training on this is in the courts,…

The very idea that LLMs are "inspired" by copyright material is so far beyond absurd I just don't know what reality you people live in. They are ingesting copyright material in order to re-use it. Yeah they remix it to add their own (incredibly annoying) tone but that's what they're doing.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#620
post #329

Earlier quoted context omitted.

Everyone on here is smart enough. Just do not participate and save your money. Do not pay for digital goods. If Netflix raises their prices, it doesn't matter because there is a torrent of all of their shows. If Spotify raises their prices, it doesn't matter because your favorite artist has their entire library in a torrent. If some game company ask you to pay real life prices for a digital costume, find the crack on…

I just can't get behind the sentiment that the unethical behavior by big companies means I get to access all the content I want for free.

Well yeah, because it’s dumb.
Post reply on HN