Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

331–340 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#331

Earlier quoted context omitted.

Everyone on here is smart enough. Just do not participate and save your money. Do not pay for digital goods. If Netflix raises their prices, it doesn't matter because there is a torrent of all of their shows. If Spotify raises their prices, it doesn't matter because your favorite artist has their entire library in a torrent. If some game company ask you to pay real life prices for a digital costume, find the crack on…

if you want to support an artist go to the show and BUY MERCH at the table! almost all of their income comes from that. the importance of buying a T-shirt at the show cannot be overstated and sometimes you get to say hi to your idol, too

It's a stupid situation, though. There are many creators I'm happy to support - but for 99% of them, I don't want their stupid merch. It's mostly low-quality garbage with high markup, that nevertheless cost something to design and produce, thus wasting both precious resources and labor - an useless tax on contributions to artists that doesn't even help anything. I really wish this wasn't necessary.

(Even the okay-quality merch is a waste, since for most artists I'd want to support, I don't identify with them enough to display that stuff, so it's again just buying to put away and eventually throw away.)

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#332
post #329

Earlier quoted context omitted.

Everyone on here is smart enough. Just do not participate and save your money. Do not pay for digital goods. If Netflix raises their prices, it doesn't matter because there is a torrent of all of their shows. If Spotify raises their prices, it doesn't matter because your favorite artist has their entire library in a torrent. If some game company ask you to pay real life prices for a digital costume, find the crack on…

I just can't get behind the sentiment that the unethical behavior by big companies means I get to access all the content I want for free.

Thank you for your service

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#333

Earlier quoted context omitted.

In Spotify’s defense, they used the pirated data only to show a proof of concept to the copyright holders, and that use was sanctioned by the local rights holders organization STIM. The copyright holders then approved their concept, and subsequently Spotify got the rights to offer their service to customers. Everybody won.

That’s not entirely true, in Spotify’s early days you could upload files to the service and listen to songs uploaded by other people. I think the majority of any song I wanted to listen to before they went Europe-only for a time was “pirated”.

Fair. I stand corrected.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#334
That they’d focus on file sharing over transformation or outputs is exactly the risk I warned the companies about in my AI report. Most datasets, like RefinedWeb and The Pile, also require sharing copyrighted workers between people who are not licensed to do that. Many works also prohibit commercial use or have patents on them.

They need to make datasets which don’t have this problem or have entities in Singapore train the foundation models within their rules. The latter has a TDM exemption that would let AI’s use much of the Internet, maybe GPL code, licensed/purchased works they digitize, etc. Very flexible.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#335
post #209

Libgen is a civilizational project that should be endorsed, not prosecuted. I hope one day people will look at it and think how stupid we were today to shun the largest collection of literary works in human history.

Libgen turns into a problem when you have a company developing generative AI with it, either giving money to GPU manufacturers or themselves with paid services (see OpenAI)

What are we actually worried about happening?

Are AI-written books getting published?

If they start out-competing humans, is that bad? According to most naysayers, they can't do anything original.

Are people asking the AI for books? And then hoping it will spit it out a human-written book word for word?

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#336

Earlier quoted context omitted.

I think they are morally required to improve the current state. - Seed the torrent and publicly promote piracy pushing lawmakers. - Contribute with digitisation and open access like Google did in the past. - Make the part of their dataset that was pirated publicly accessible. - Fight stupid copyright laws. I can't believe that copyright lasts more than 20 years. No field moves that slowly, and there should be tighter…

Copyright and patent aren't the same thing. "Fast moving field" doesn't make sense in terms of copyrights. There's no reason the copywriter should last some minimum duration after the life of the creator. If I write a really popular book, I don't want Hollywood to make it into a movie without compensating me just because they waited a few years

Fast moving field does make sense in terms of copyright because the knowledge is recorded in documents which are then copyrighted. E.g. research papers.

> If I write a really popular book, I don't want Hollywood to make it into a movie without compensating me just because they waited a few years

I genuinely don't understand this. Even at a decade copyright, pretty much anybody who was going to buy the book and read it has already done so. It costs you virtually nothing in sales, and society benefits from the resulting movie.

Your goal is to deprive everyone of having a movie, because someone who isn't you is going to make some money that was never going to you anyways? Your goals for copyright appear to be a net negative to the system that enforces copyright, which begs the question why should the system offer protection at all?

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#338

Earlier quoted context omitted.

The hotel and taxi industry were legit terrible before those two disrupted them. Laws are ment to be broken. Especially in cronist systems where incumbents write the laws.

Hotels were just fine. Taxis were discriminatory and "uncool" to the point were Uber has saved thousands by preventing drunk driving. Now if you go out with the boys and get drunk, it's a 30 second casual call to get an Uber and get home. Live in a neighborhood Taxis are afraid to service,you can either make some extra income working for Uber or use it yourself. When Ubers used as its intended purpose, to basically m…

[deleted]

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#339
post #123

Earlier quoted context omitted.

I think you’re overstating its importance. The internet already makes it possible to order almost any book in existence and have it arrive at your doorstep within a week or so, or often on your ebook reader instantly. And your local library probably participates in an interlibrary loan system that lets you request any book held by any library in the country for free . LibGen gives you access to a much smaller body of…

[flagged]

I definitely do know what I think.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#340
post #166

Earlier quoted context omitted.

Airbnb and Uber have showed us that laws matter only to the extent that the political will to enforce them exists. Throw enough lawyers and lobbying money at the problem and the laws can simply be re-written to be friendlier to your business model.

The reason there was no political will to punish Airbnb and Uber for violating the law was that initially they were subsidized with VC money and so were able to undercut traditional hotels and taxis on price. In the world of tradable goods, pricing below cost with the intent of putting competition out of business so you can raise prices later is known as "dumping" and is itself illegal.

I rooted for Uber to smash the Taxi cartels. Let us not forget that Taxi Cartels were also insidious beasts. Taxi drivers abused their walled garden with their price gouging by taking longer routes, refusal to take a credit card, and extremely poorly maintained fleets of vehicles. I have had mostly good experiences with Uber, whereas I had experiences that mostly bordered on general condescension toward me whenever I took a ride in a Taxi. I am glad the political will to block Uber never materialized.
Post reply on HN