Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

111–120 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#112
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

> Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same.

I don't understand why you wouldn't just buy copies of the books. Seems like such a relatively inexpensive way to strengthen your legal case.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#113
post #89

Earlier quoted context omitted.

The legal system is built to favor large corps and capital owners. See Katharina Pistor books for instance.

I think it’s the other way around. Those large entities break all the same laws and rules as others and then get to the point where they can influence the creation of a regulatory moat around themselves to prevent competitors from taking the same path as them.

True but lets take examples one by one to see what we can learn : Spotify was doing illegal things until they made a deal to become legal and not to be trialed over what they done. Seems like business deals is what saved them, not regulatory capture (the regulations around IP for music pre existed Spotify)

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#114

Earlier quoted context omitted.

No we just need to enforce the existing laws. And the legal system is for humans not computers.

The existing laws are a problem, and are not enforced in a fair and just manner. Yes, the legal system is for humans, but we can use technology to improve the system for humans, so it's faster, better and more fair, because humans aren't perfect, and now we have technology to be better than the system create a long time ago. You don't think the legal system should run on pens and paper right? Adapting to typewriters,…

Aren't all LLMs based on models published by the big two or three, so built on IP theft and if you're using them you are guilty of handling stolen goods?

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#115
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

" If you plug a laptop into a closet at MIT to download some scientific papers you forfeit your life." This is exactly what I immediately thought while reading the article. It almost feels like the legal system only punishes general public, while most of these guys are above it.

It doesn't "seem". The entire system in most countries works, by design, that way because the people in power trade in influence at a different plane.

That's why democracy often feels "failed" in that no change can be achieved because "it's just more of the same". Few Lobbyists representing the interests of a few people have more power than millions voting differently.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#116

Earlier quoted context omitted.

Library genesis

Weird shenanigans are happening in libgen at the moment; better go through Anna's Archive to look for the items you want, it will link you to the corresponding mirrors more reliably. At least this has been the recent experience of a friend who used libgen and anna's archive to download legal, public domain works!

No, AA is rate limited to being unusable, while libgen is fast enough.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#117
post #98
post #91

Earlier quoted context omitted.

Meta is not “innocent”, and comparing this instance with Swartz is a huge offense to his legacy.

I don't think you've read the parent comment correctly?

Parent comment implies Swartz was guilty of some degree. I vehemently disagree with that.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#118
We have at least 4 types of ill-defined concepts of property in the 21st century , largely due to our laziness, intellectual inertia and lack of motivation to make forward-thinking definitions for the coming age of AI and ubiquitous access to all information and all communication.

1) the concept of copyright is as old as the word suggests (copies are the least of our worries going forward - it should be possible to define processes for exploitation of ideas in a fair way)

2) we allow humans to learn from other people's ideas and transform them to commercial products and the same should happen for AIs in the future

3) we have an ill-defined concept of "personally identifying information" which gives people ownership to information that others have created via their own means - there should be better ways to ensure a level of privacy (but not absolute privacy) without overly-broad, nonsensical definitions of what is personally protected information

4) We allow social media and other telecommunications media to arbitrarily censor people's speech without recourse. This turns people's speech to property of the social media companies and imposes absolute power on it. This makes zero sense and is abusive towards the public at large. We need legal protections of speech in all media, not just state-owned media.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#119
post #20

A good chance for federal prosectutors to "send a message" as they did with Aaron Swartz but I don't see things going that way.

If you were wondering why meta was making a lot of donations to the new government (including settling a lawsuit for 25 million with the New president, 1 million to the inauguration)…. I suspect there will be no federeal charges.

The rules have always seemed different for corporations regardless.

https://www.businessinsider.com/trump-settles-lawsuit-meta-m...

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#120

Earlier quoted context omitted.

the english empire once tried to mantain a monopoly over steam loom machines the americans cheated their way to competition, heck, even before that, the english empire got jumpstarted by stealing gold from the spanish (who were themselves exploiting it away from aztec and other mexican natives) I'm saying it's business as usual, but also, culture doesn't work like tangible physical widgets so we must stop letting a f…

People criming in the past is not an excuse for companies committing crimes today. You’re excusing lawlessness. Cain killed Abel and got away with it!! I can kill someone today too!!!

I think it’s fine to criticize the hypocrisy of viciously defending the copyrights you own, while gleefully running roughshod over the ones you don’t.

But it’s also possible that copyright as a concept, or in its current implementation, is bad and unjust.

I’m sure some copyright holders would like nothing more than to see an argument that elevates copyright violation to the level of murder, morally or legally. But I think it’s more akin to jaywalking - violating an unjust law that mostly shouldn’t exist.

Post reply on HN