Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

921–930 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#921
post #857

Earlier quoted context omitted.

Is there any other possible outcome than a fine? That too one which will not really affect Meta's overall earnings

> Is there any other possible outcome than a fine? Yes, of course. It's quite possible that judges realize that if they restrict training data to licensed materials, LLMs will become stupid and China will overtake the US to become the leader in AI, and because that can't happen , they'll make up some reason to make training on unlicensed data legal. It's definitely fair use! I'm not even joking. Last time the US Supr…

> Last time the US Supreme Court basically said "Android is too important, we have to declare its use of Java API fair use."

Yes, I remember that case. I do think a hefty fine can be sufficient here though. Meta has been paying out court ordered fines for years, none of which dent their growth much. And current models are trained on those models + datasets are scrubbed more vigorously now so future rehashes of this shouldn't be an issue. Also, most data online has already been hoovered up, according to Ilya the former Chief Scientist of OpenAI, so this will likely not pop back up again.

But good point, courts are often more practical given the law than we perceive them to be.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#922
post #913

Earlier quoted context omitted.

Imo it’s not about you accessing things you want for free. If your family purchased a disc copy of the goonies before you were born and you watched it as a kid, your accessing of that content you wanted for free has no moral bearing. The core question is what impact does your consumption have, and I don’t think that participating in the streaming landscape is making things any better for anyone but their ceos.

The comment I was responding to certainly seemed to be advocating uninhibited free access via torrents, but your point is reasonable. I don't think streaming is a great way to support artists either, though I do think it's better than nothing.

I just don’t think streaming is really working for anybody. Artists are being given less and less creative control as time goes on as each streaming company attempts to optimize itself in a largely fixed market. It’s working the same way cable used to work. I think a disruption is in order, but last time disruption started with people moving to nascent torrent sites.

I guess the way I’d put it is that if you can only get some particular show through one company, that company gets to treat you shitty, cause where else are you gonna go. Torrenting, even just widespread knowledge of torrenting, gives the customer more leverage.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#924
post #25

ebooks are a 1-2 mb each max. 81.7 TB are a lot of books, like 42-85 million books.

The article says they got datasets from Anna's Archive. It was most likely the scihub/libgen torrent which is 96.0 TB right now and contains 92,872,581 files. That's about 1 megabyte per file. https://annas-archive.org/datasets

Where does one find these torrent datasets? Did they download the books in bits and pieces or as a single huge multi-TB file?

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#925
post #802

Earlier quoted context omitted.

I agree that removing bad laws is good. I think by introducing the second, culturally charged topic (1.) taxi cartels, 2.) recreational drugs) you diminish the possible interpretations of your perspective. The other downstream conclusions make sense too, but the linkage is more opaque making it difficult to appreciate. Also hard to acknowledge is--who decides which laws are "bad"? Generally, societal outcomes should…

> who decides which laws are "bad"? It's better to ask the question in a different way. We know what bad laws are . They're laws that benefit some interest group at the expense of the general public, e.g. by constraining competition or diverting tax dollars to cronies. So the question is, how do you eliminate bad laws? This isn't a question of what a hypothetical legislature should do if it was full of good faith act…

That makes sense but seems like it would only actually be a subset of bad laws. I mainly mean to highlight that it's not a comprehensive way to identify bad law.

> how to structurally align the incentives of a real legislature with the interests of the general public

This seems like a critical nuance that, like you said, needs a structural solution. I have no actual idea, but conceptually this seems like it would eliminate a subset of particularly bad laws and actions (e.g. members of the legislature trading on their insider information) which have outsized, negative outcomes for the public. But we also rely on that very rule making body to essentially self-govern. And such a grass-roots movement of reforms to put the public first seem unlikely given the attitudes and sensationalizing behaviors present in the members of that body.

I avoid politics because of just how disaffecting it is to think about most of these details.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#926

Earlier quoted context omitted.

Whether that falls under fair use is highly debatable.

It's going through the courts right now. We'll probably have an answer in a year or two.

I felt for a long time that it should be fair use. If an LLM can abstract what it learns from the copyrighted work, then that seems "fair" because that's what humans do.

But ... as I've thought about it more, it doesn't really feel just to me. The kind of value reaped from the works seems to suggest that the creator is due some portion of that value. Also, in practice - there's just an absolutely enormous amount of knowledge that can be consumed from the public domain. Even if Meta, OpenAI and friends decided to license a ~small handful of the long-term archives of some globally-read newspapers, they could get very broad and deep knowledge about the events, trends, terms of the last century to fill in a lot of gaps.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#927

Earlier quoted context omitted.

It isn't that someone was hurt. We have one private entity gaining power by centralizing knowledge (which they never contributed to) and making people pay for regurgitating the distilled knowledge, for profit. Few entities can do that (I can't). Most people are forced to work for companies that sell their work to the higher bidder (which are the very entities mentioned above), or ask them to use AI (under the conditi…

Then you should be in support of OSS models over private entity ones like OpenAI's.

Like supporting Android Open Source Project… until Google decides to move the critical parts into Google Play Services? I run GrapheneOS (love it) but almost no banks will allow non-Google-sponsored Android ROMs and variants to do NFC transactions because… the AOSP is designed to miss what Google actually needs.

Idem with ML Kit loaded by Play Services, which makes Android apps fail in many cases.

And I'm not talking about biases introduced by private entities that open source their models but pursue their own goals (e.g geopolitical).

As long as AI is designed and led by huuge private entities, you'll have a hard time benefiting from it without supporting the entities' very own goals.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#928
post #913

Earlier quoted context omitted.

The comment I was responding to certainly seemed to be advocating uninhibited free access via torrents, but your point is reasonable. I don't think streaming is a great way to support artists either, though I do think it's better than nothing.

I just don’t think streaming is really working for anybody. Artists are being given less and less creative control as time goes on as each streaming company attempts to optimize itself in a largely fixed market. It’s working the same way cable used to work. I think a disruption is in order, but last time disruption started with people moving to nascent torrent sites. I guess the way I’d put it is that if you can only…

But is torrenting better than buying music e.g. through bandcamp or directly from the artist's label? I do a combination of streaming for exploratory listening and buying to support artists. And live shows when I can. I'm not trying to suggest I'm holier than thou, but I really don't think torrents are the answer to this equation. However I strongly agree that the current system isn't working for artists and needs disruption.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#929

Earlier quoted context omitted.

I honestly suspect fairly little would change. The US operates with a 20 year copyright for nearly 200 years, these long copyrights are actually far newer. Also, you are not owed a monopoly on arrangement of words enforced by the public. There are plenty of other places to spend tax dollars.

What tax dollars are we spending enforcing copywrite laws?

Who do you think pays for prosecutors, lawyers, jails, and investigators going after pirating sites? DMCA claims are worth less than the paper they're printed on if not for the threat of all that.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#930

Earlier quoted context omitted.

We're sick of the double standards. https://en.wikipedia.org/wiki/Aaron_Swartz#United_States_v._... https://en.wikipedia.org/wiki/Aaron_Swartz#Death While Aaron Swartz was bullied to suicide, these corporations will walk free and make billions. I say give every tech CEO the Swartz treatment, then change the law.

Why not change the law first? Two wrongs don't make a right. If a law is unjust, then what good is there in continuing to punish people who have broken it, just because other people have been punished in the past? Either you think the law is just or unjust. If you think it's unjust, I don't possibly see how you think people should be punished for it. Meta wasn't responsible for what happened to Aaron Swartz.

Motivation, consequences and fairness of how rules are applied are all relevant to the ethics of a particular action.

Meta is systematically abusing copyright law for personal profit, appropriating the labor of hundreds of thousands of authors to line their own pockets without contributing anything back. That free-riding is anti-social behavior, a betrayal of the social contract. A society that permits that without censure isn't going to keep having people create for very long: who is going to create just so billionaires can get richer, with nothing in it for them?

Especially not when the same billionaires sue people for violating copyright of _their_ software. Hypocrisy in service of exploitation and greed is especially noxious.

At the same time, copyright laws are written to benefit those billionaires, keeping our mythology in private hands long past our lifetimes. Put copyright back to thirty years, strip all business methods and software patents (which should never have been a thing in the first place): then Meta will have plenty of content for their LLMs _and_ it's software will start coming out of copyright ten years from now, making an actual contribution to human knowledge instead of just pillaging it.

The central tenet of conservativism is that there is a group the law protects but does not bind and a group the law binds but does not protect. This is what Meta is doing.

They could have gone and lobbied to loosen copyright legislation. Heck, they could have gone through a DMCA exception process, which doesn't even take a new law. Instead, they figure they are powerful enough that society doesn't apply to them.

If we let them, we're suckers.

Post reply on HN