Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

531–540 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#531

Earlier quoted context omitted.

From the article, they took steps to avoid using their IP addresses. Individuals doing the same using a VPN are pretty much immune from any legal issues.

This is one small blip in an incredibly long history of companies being able to not care about copyright while individuals must. Workarounds with a VPN are great and all, but they are a band-aid on a systemic problem. (You are not immune , by the way, if your VPN company is subject to a subpoena and isn't one of very few actually no-log services)

IME companies take copyright way more seriously than individuals. e.g. my last 2 jobs have had scanners to ensure we're not accidentally pulling in GPL code to our products, and one of those was a startup. I'd be surprised if corporate security software weren't looking for torrent clients and if you wouldn't get fired for torrenting on corporate machines or networks at most companies. Meanwhile the same people setting those security policies have a 100TB array at home with fully automated pirating setups. They very much don't care personally, but it's a huge business risk.

In high school/university in the 00s, everyone casually pirated things. In college people passed around a USB drive with all of the books needed for our degree program. People in the dorms traded music collections with 10s of thousands of songs. Tellingly, Apple advertised that iPods could store 10,000 songs, which approximately zero people could afford to buy legitimately. If anything, the consequences for piracy have gone down since then, but streaming is convenient enough and phone storage/UX is hobbled enough that people pay.

In any case, I think the other poster is right that companies flouting copyright law is a good thing. It stops us from pretending that it's helpful for the little guy, making it easier to argue for abolition or vastly reducing the length. That they did it to build an open model is even better: it shows directly the kinds of benefits copyright is taking from us. We should be looking to scan every book out there to build better training sets (and better indexed search into scholastic datasets; at this point all of Anna's Archive only costs a little over $11k in raw storage, which puts it into "affordable as an upper middle class home library" territory. In another few years, it may be affordable to nearly everyone. Better ML models could help here with better compression as well), but copyright law restricts use of works dating back to a time before electrification was widespread. Obviously they're an evil company in general, but llama was an actual good deed from them.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#532

Earlier quoted context omitted.

I think they would need to have some explicit contract every time they want to sell the book then, though. I don’t think I am bound by some random terms someone writes into a book I’m buying. Those probably are only binding if a reasonable person would notice them before sale.

If you arrive at the point of being able to buy that book, it means it has passed the publisher's hands and I would think, that the publisher was OK with those terms then, and limiting the usage of the text may in fact be effective. If it was self-published, then even more so.

But the license restriction would have to apply both to the publisher and the customer.

If I go to the bookstore, buy the book, make a scan, and train an LLM with it, how would you enforce your license as an author? The customer never knew that he shouldn’t have been allowed to train LLMs.

Edit: I think I misunderstood the original comment, I thought the idea was to sell books and restrict use for LLM training. If we’re only talking about stuff that’s publicly released, the restriction should be possible.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#533

Earlier quoted context omitted.

The reason there was no political will to punish Airbnb and Uber for violating the law was that initially they were subsidized with VC money and so were able to undercut traditional hotels and taxis on price. In the world of tradable goods, pricing below cost with the intent of putting competition out of business so you can raise prices later is known as "dumping" and is itself illegal.

I rooted for Uber to smash the Taxi cartels. Let us not forget that Taxi Cartels were also insidious beasts. Taxi drivers abused their walled garden with their price gouging by taking longer routes, refusal to take a credit card, and extremely poorly maintained fleets of vehicles. I have had mostly good experiences with Uber, whereas I had experiences that mostly bordered on general condescension toward me whenever I…

Political will shouldn't be required to enforce existing law. If i started a 1 man illegal taxi service it would be shut down even though it has little effect on the community, but saudi vc funded startup wasn't shut down even though it violated laws in every major city. That is a weird asymmetry as a us citizen.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#534

Libgen is a civilizational project that should be endorsed, not prosecuted. I hope one day people will look at it and think how stupid we were today to shun the largest collection of literary works in human history.

I wonder how much more libgen traffic can be attributed to the lawsuit.

When Metallica sued Napster, for many people the reaction was, "wait I can download music for free?"

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#535
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

> If you plug a laptop into a closet at MIT to download some scientific papers you forfeit your life. In case anybody here doesn't know, that's a reference to Aaron Swartz, an activist (and Reddit co-founder) that was risking 35 years in prison and a $1 million fine just for downloading a lot of academic papers from JSTOR. He eventually took his life because of the pressure. May his soul rest in peace.

Thanks, I thought it was a sarcastic reference to torrents. So this cleared that up

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#536

I don't understand why it's even a question that Meta trained their LLM on copyrighted material. They say so in their paper! Quoting from their LLaMMa paper [Touvron et al., 2023]: > We include two book corpora in our training dataset: the Gutenberg Project, [...], and the Books3 section of ThePile (Gao et al., 2020), a publicly available dataset for training large language models. Following that reference: > Books3…

There are two different things when it comes to discussing training LLM's on "copyright" protected data, and I almost never see people differentiate. 1.) Training on copyright that is publicly available. You write a poem and publish it online for the world to read. That is your IP, no one else can take it an sell it, but they are free to read and be inspired by it. The legalitly of training on this is in the courts,…

I never gave my poem to Facebook. My site is for humans. And there was absolutely no problem with that website being public, until Facebook et al wanted to move the goalpost.. again. Remember when companies started to claim that their abuse is on you, because you failed to publish the correct headers/robots.txt and their bot needs to be told the rules in specific language? And now we get the same attempt at making such distinction again, just this time its our fault for .. having a public website in the first place (should have operated a paywall, duh!)

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#537

Earlier quoted context omitted.

We're sick of the double standards. https://en.wikipedia.org/wiki/Aaron_Swartz#United_States_v._... https://en.wikipedia.org/wiki/Aaron_Swartz#Death While Aaron Swartz was bullied to suicide, these corporations will walk free and make billions. I say give every tech CEO the Swartz treatment, then change the law.

The lesson here is make sure you only break the rules in the limits of severity that your wealth class allows. MIT students will get away with breaking bigger rules than community college students will.

Ah, you're citing that inviolable document, the United States Constitution, which brought forth the even-handed dawn of a legal caste system.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#538

Earlier quoted context omitted.

Everyone on here is smart enough. Just do not participate and save your money. Do not pay for digital goods. If Netflix raises their prices, it doesn't matter because there is a torrent of all of their shows. If Spotify raises their prices, it doesn't matter because your favorite artist has their entire library in a torrent. If some game company ask you to pay real life prices for a digital costume, find the crack on…

if you want to support an artist go to the show and BUY MERCH at the table! almost all of their income comes from that. the importance of buying a T-shirt at the show cannot be overstated and sometimes you get to say hi to your idol, too

A lot of artists are under a 360 deal and they take a cut of everything, so that might not always be true.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#540

Earlier quoted context omitted.

If you arrive at the point of being able to buy that book, it means it has passed the publisher's hands and I would think, that the publisher was OK with those terms then, and limiting the usage of the text may in fact be effective. If it was self-published, then even more so.

But the license restriction would have to apply both to the publisher and the customer. If I go to the bookstore, buy the book, make a scan, and train an LLM with it, how would you enforce your license as an author? The customer never knew that he shouldn’t have been allowed to train LLMs. Edit: I think I misunderstood the original comment, I thought the idea was to sell books and restrict use for LLM training. If we…

Whether you make a scan of it or not, the license applies to the IP, I guess (IANAL).

Whether the shop makes a scan should not affect you as the buyer of the actual book. What does the scan have to do with you?

Whether the author learns about that scan and perhaps training of some LLM using the scan or not, does not change the legality of it.

Post reply on HN