Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

771–780 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#771
post #329

Earlier quoted context omitted.

Everyone on here is smart enough. Just do not participate and save your money. Do not pay for digital goods. If Netflix raises their prices, it doesn't matter because there is a torrent of all of their shows. If Spotify raises their prices, it doesn't matter because your favorite artist has their entire library in a torrent. If some game company ask you to pay real life prices for a digital costume, find the crack on…

I just can't get behind the sentiment that the unethical behavior by big companies means I get to access all the content I want for free.

I think the point is, using Spotify is already essentially listening to your artists without supporting them. You might as well do that without supporting a company that is harming them.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#772

Earlier quoted context omitted.

Yes. And the problem here isn't that companies get away with doing things like this, the problem is that individuals don't. Attempting to lock information behind a nightmarish legal system is the problem. I'm pretty much at the point now where I don't buy the "copyright incentivizes creation" argument any more. Copyright, like advertising, incentivizes creation by enormous corporations, but also like advertising it i…

Sure creative people will always create but the scope of that creativity will be limited if we do away with intellectual property. Steve Spielberg would probably always have created movies, but he wouldn't have been able to make Jurassic Park, Saving Private Ryan,or Indiana Jones without capital from the studio system, and the studio system wouldn't have provided him with that capital of they couldn't extract economi…

Do we need to always have big-budget films and productions? Perhaps we should live smaller, and enjoy local art and low-budget films. Do I really care that Jurassic Park was made? I could read the book and it's more detailed and imaginative anyways, and any lessons to be learned are definitely better when read than when watching a blockbuster CGI film with more effects than message.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#773
post #20

A good chance for federal prosectutors to "send a message" as they did with Aaron Swartz but I don't see things going that way.

Well of course, bullies always prefer targets that can't fight back. That itself is unfortunately a basis of the legal system from it being run on flawed monkey brains. Why else is hitting vulnerable children okay but getting into a consensual bar fight illegal?

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#774

It really makes you think about those crazy internet folks from back in the day who thought copyright law was too strict and that restricting humanity to knowledge in such a way was holding us all back for the benefit of a tiny few.

Probably the single biggest thing I learned growing up is that you can safely live by "Everyone is in it for themselves". It's incredibly rare to find people who hold ideals that are detrimental to their own life.

Yep, you even see it on HN with artists and devs complaining about AI, especially when things like ChatGPT and Stable Diffusion were first announced. People who were pretty lax about copyright when it didn't affect them personally suddenly became copyright maximalists, talking about "stealing, theft, etc" Since then, people have calmed down and realized that AI is simply a tool like any other.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#775
post #740

Something tells me uncle Donald will exonerate his new favourite lapdog from any criminal or civil liability.

IANAL but the pardon power (A) only extends to criminal punishments, not civil liabilities and (B) copyright lawsuits can be launched by anybody, not just the Department of Justice. So, barring further Might Makes Right shit--which I'm not willing to fully rule out--Trump can't fully shield Zuckerberg et al.

They can also sue in France, or Spain, or Japan.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#776

The more I learn about how AI companies trained their models, the more obvious it is that the rest of us are just suckers. We're out here assuming that laws matter, that we should never misrepresent or hide what we're doing for our work, that we should honor our own terms of use and the terms of use of other sites/products, that if we register for a website or piece of content we should always use our work email addr…

>What we should have been doing all along is YOLO-ing everything

No it isn't. The actual sucker attitude is copying what they do. You should act morally and with integrity out of respect for yourself. I never had any illusions that large tech companies act with respect towards the law, but it also has nothing to do with me.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#777

This reminds me of Peter Sunde's "komimashin" https://www.engadget.com/2015-12-21-peter-sunde-kopimashin.h... It's obviously absurd to enforce copyright as bytes are copied around instead of as it is used. Training an LLM is a different thing than re-hosting and giving away copies to other people. If you don't want people to transform your works - keep them private. You don't own ideas.

As the article says, Meta /was/ giving away copies to other people by seeding the libgen torrents. This isn't the usual case of "should companies be allowed to train on books".

Then it's a simple case of a rights holder taking them to court.

What's the fuss about LLM training in this thread then?

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#778
post #715

The more I learn about how AI companies trained their models, the more obvious it is that the rest of us are just suckers. We're out here assuming that laws matter, that we should never misrepresent or hide what we're doing for our work, that we should honor our own terms of use and the terms of use of other sites/products, that if we register for a website or piece of content we should always use our work email addr…

This sort of mindset is devoid of morals and honor. Don’t fall into the this mindset trap. Like when Trump said he is “smart” for evading taxes during the presidential debates (IIRC the first ones, not recent ones). It’s absolutely despicable. Have a moral compass. Treat people fairly. Be nice. Let’s be better than toddlers who haven’t learned yet that hitting is bad, and you shouldn’t do it even if mommy and daddy a…

I agree with you and will die to defend that position, but what existential reasons do we have to behave well? In reality, it seems like humans are just a bunch of animals, and the only important thing is survival.

My wife, just today, told me that she was very upset that I refused to interview and take jobs for things like building weapons, the panopticon, or advertising (two of those are the same thing), which I refuse to do because of my personal morals and ethics. How do I explain to her that I just can't do that, and give her a good reason why we should lose our home and live in one room with her mother because of my brain refusing to work in such industries? I really want to know, so I can explain to her and my son why such things matter, because for some reason they are concrete and foundational in my brain, there is no changing that.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#779

We need better laws that would create a better way to do this legally whilst compensating rights holders.

I really don't think that Meta did this because the alternative would have been too onerous; they are a huge org, they could work through whatever loopholes required. They did it because it would have cost money and there will be no penalty for not paying.

So, if they're sued in Japan, or France, do you think that the courts will take any special measures because it's a valuable American corporation?

I suspect that if the case is reasonable they will just convict, and quickly-- appeal denied and all simply because the laws are so straightforward.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#780
post #685

Earlier quoted context omitted.

> getting a taxi license was a serious monetary investment. People took our huge loans for this it was a terrible system that sucked for everyone involved. For all of Uber's flaws, would you rather go back to that today? really??

> would you rather go back to that today? really?? I absolutely said said no such thing. There are good ways to change things and bad ways to change things. Allowing a private entity reap huge profits by blatantly breaking rules and screwing people is not a good way to change things.

There was no other way to change things on less than a generational timescale.

If governments don't like it, well, bummer. They were supposed to serve the people, not the incumbent taxi cartels. They failed, so "we the people" routed around them.

Post reply on HN