Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

311–320 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#311
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

Some can pirate on a large scale and see no repercussions.

Some can steal from stores and see no repercussions.

Some can steal from others and see no repercussions.

Some can violently harm others and see no repercussions.

Some can damage property and see no repercussions.

Some can’t. This world is not right.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#312
post #226
post #218

Earlier quoted context omitted.

Perhaps they just did, or we are doing it - basically this should lead to abolition of copyright to any published article there is. Not sure how’d it impact open source, we’ll either have all of it open, or none at all.

Even without copyright there are trade secrets, not to mention trademarks and patents. Maybe we could get rid of the latter, but I think we’d need to be pretty heavily into socialist utopia before considering nixing the former two!

Trademarks and patents are very different from copyright. Trademarks especially so because they aren't designed to "own" knowledge, just to prevent confusion about who made a product or what it is.

"Intellectual property" is an abomination of a term because it conflates 3 separate mechanisms with differing goals, pretending that they're related in any meaningful sense.

Patents protect a process. Trademarks protect identity. Copyright protects knowledge. Disparate mechanisms for disparate goals.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#313

Earlier quoted context omitted.

" If you plug a laptop into a closet at MIT to download some scientific papers you forfeit your life." This is exactly what I immediately thought while reading the article. It almost feels like the legal system only punishes general public, while most of these guys are above it.

It's an oligarchy, always has been. I don't know how colossal the pile of evidence supporting this has to get before people finally accept it.

They are paid, handsomely, by it. Or otherwise brainwashed by it. And pummelled into ignorance by it, as they are told that to understand is stupid or delusional, knowledge ends at STEM, and the world only exists for efficient production of capital products.

The poets write laments about such false ages. Prophecies were written about such ages thousands of years ago.

The cycles are larger than us all.

One stable insight is that the chaos breeds possibility, and thus hope. In the meantime, however…

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#314
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

[flagged]

I don't think it's right to downplay the disproportionate response the FBI had to Aaron's actions. He was initially being threatened with 50 years in prison and a $1 million fine, the stress from which sent his mental health spiraling and in no small way contributed to his suicide. I think the original point of the person you are responding to still stands.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#315
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

> once people started uploading copyrighted TV shows to it

End users, not YouTube employees, right? And they would take things down following DMCA requests and what not, right? So, pretty much following the law?

> Google itself got big by indexing other people's data without compensation

Scraping public websites to build a search index isn't the same as making LLMs that can recreate the source verbatim devoid of even attribution. I do agree there's an argument to be had about the LLM's transformative nature in the end though.

> Spotify's music library was also pirated in the early days

Not any version generally available to the public, and with the copyright holder's permission to do so.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#316
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

And Hollywood was created on the west coast because for intellectual property it was still the far west and it allowed them to ignore patents on movie technologies.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#317
post #263
post #241

Earlier quoted context omitted.

…why? Will people buy less books because we have intuitive algorithms trained on old books? Personally, I strongly believe that the aesthetic skills of humanity are one of our most advanced faculties — we are nowhere close to replacing them with fully-automated output, AGI or no.

old books? i can imagine the shit/hallucinated-like generative AI we would have if the training weight was restricted to public domain stuff... i think when chatGPT was around version 2 or 3, i had extracted almost 2 pages (without any alteration from the original) with questions that considered the author from this book here, https://www.amazon.com/Loneliness-Human-Nature-Social-Connec... now it's up to you to think…

The answer is to censor the model output, not the training input. A dumb filter using 20 year old technology can easily stop LLM's from verbatim copyright output.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#318
post #305

> By September 2023, Bashlykov had seemingly dropped the emojis, consulting the legal team directly and emphasizing in an email that "using torrents would entail ‘seeding’ the files—i.e., sharing the content outside, this could be legally not OK." I'm pretty sure you can theoretically download torrents without seeding, although this is frowned upon. If they really seeded (with full bandwidth?) that's indeed pretty br…

You can, in transmission for example you can just set the seed percentage to 0%. I recognise that this makes me a bad torrenter, but I've been told in the past that my ISP wont be too happy about me seeding, and they already do something screwy to torrents I access through the surface web, so I'm just playing it safe

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#319
So they're gonna go through every book that was stolen and apply the appropriate penalty, right? Each copyrighted work has a minimum penalty of $750 under the DMCA. That will be applied fairly in order to ensure that the rights holder is made whole by the infringer, right?

It's so funny to see the law blatantly ignored by the overlords. Like, there isn't even a pretext anymore. They just steal what they want and budget for the fines and campaign donations to make the consequences go away.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#320

Earlier quoted context omitted.

The reason there was no political will to punish Airbnb and Uber for violating the law was that initially they were subsidized with VC money and so were able to undercut traditional hotels and taxis on price. In the world of tradable goods, pricing below cost with the intent of putting competition out of business so you can raise prices later is known as "dumping" and is itself illegal.

The traditional taxi industry was rife with corruption, bad experiences, and poor service in many jurisdictions before uber/lyft. As terrible of a human being that I think Travis Kalanick is, it was only going to take lawbreaking to overcome such a tainted system. Medallion systems often prevented any competition, sometimes to absurd effect. The number of licenced taxis often didn't keep pace with population growth,…

It sucked, but not everywhere equally. Meanwhile, Uber rode their one-trick pony (an app), which everyone quickly replicated, all the way to upending taxi businesses worldwide, thanks to their infinite money supply letting them survive long enough in any new market to get the public behind them, which took away support from local regulators trying to keep the market from being gutted by what at this point was a multinational corporation (and technically a criminal enterprise).

Sure, taxi services aren't usually known to be paragons of virtue, but then they weren't that bad everywhere; Uber is just another case of an US org trying to address an US-specific problem and then bludgeoning the entire world with their solution, whether the rest of the planet has such problems or not.

Post reply on HN