Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

221–230 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#221
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

" If you plug a laptop into a closet at MIT to download some scientific papers you forfeit your life." This is exactly what I immediately thought while reading the article. It almost feels like the legal system only punishes general public, while most of these guys are above it.

Welcome to the two-tier legal system of the modern world. Why obey the law when the penalty is a rounding error?

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#222

I strongly urge people to read Thomas Babington Macaulay's speeches on copyright, its aims, terms, and hazards. Very well reasoned and explained. In particular, people often cited the case of authors who had died leaving a family in destitution, and claimed that copyright extension would be a fair way of preventing this, but in most cases the remaining family had never held the copyright; the author had initally sold…

I’m a huge IP hater and am sure that happens, but to be fair, letting copyright extend past death also increases the amount the author can sell it for in the first place.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#223
post #118

We have at least 4 types of ill-defined concepts of property in the 21st century , largely due to our laziness, intellectual inertia and lack of motivation to make forward-thinking definitions for the coming age of AI and ubiquitous access to all information and all communication. 1) the concept of copyright is as old as the word suggests (copies are the least of our worries going forward - it should be possible to d…

>we have an ill-defined concept of "personally identifying information" which gives people ownership to information that others have created via their own means - there should be better ways to ensure a level of privacy (but not absolute privacy) without overly-broad, nonsensical definitions of what is personally protected information

What information about me could a corporation create via its own means that would be legally protected but shouldn't be? PII is generally information that a corporation collects. Unless you mean that my cellphone provider creates the association between my name and phone number and should therefore be able to do with it as they please?

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#224
post #215

Earlier quoted context omitted.

We don’t allow indiscriminate human experimentation in medicine. We have crimes against this, and yet we still have new medicines. Sure, it won’t be as quick if we could just use humans as test subjects from the start, but that’s an unethical line. Innovation done immorally is progress that shouldn’t have been made. The ends don’t justify the means, but I’m not an ethical nihilist. The crime is downloading and copyin…

Those medical policies have condemned thousands, possibly millions, to lives of unnecessary pain and suffering. They're more damaging than copyright.

Ok, I’ll go tell the Nazis that their medical experiments using live humans were A-OK!!!

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#225

It really makes you think about those crazy internet folks from back in the day who thought copyright law was too strict and that restricting humanity to knowledge in such a way was holding us all back for the benefit of a tiny few.

> crazy internet folks from back in the day

You mean Electronic Frontier Foundation? https://www.eff.org/issues/innovation

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#226
post #218
post #185

Earlier quoted context omitted.

You're conflating different problems. Big corporations are too big, they should just not exist. When you have corporations more powerful than the government of the biggest states, it's a bug, not a feature. The IP laws may need rethinking. Saying that they should disappear because big corporations are above the law doesn't help, though. First kill the big corporations, then think about fair laws. Changing the law now…

Perhaps they just did, or we are doing it - basically this should lead to abolition of copyright to any published article there is. Not sure how’d it impact open source, we’ll either have all of it open, or none at all.

Even without copyright there are trade secrets, not to mention trademarks and patents. Maybe we could get rid of the latter, but I think we’d need to be pretty heavily into socialist utopia before considering nixing the former two!

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#227

Earlier quoted context omitted.

People criming in the past is not an excuse for companies committing crimes today. You’re excusing lawlessness. Cain killed Abel and got away with it!! I can kill someone today too!!!

I think it’s fine to criticize the hypocrisy of viciously defending the copyrights you own, while gleefully running roughshod over the ones you don’t. But it’s also possible that copyright as a concept, or in its current implementation, is bad and unjust. I’m sure some copyright holders would like nothing more than to see an argument that elevates copyright violation to the level of murder, morally or legally. But I…

the reform needs to happen at the layer where whether a copyright is valid or not is decided upon, not before (at the point of "should copyright exist") and not after (enforcement).

a world without copyright means those with the largest advertising budgets will reap nearly all the rewards from new IP created by small artists. BigCorp Inc. can just sit around and wait for talented musicians to post something interesting on soundcloud, for example, then just have their in-house people copy it and push it out to radio and streaming platforms via their massive ad budgets and favorable relationships for getting new material onto the waves immediately. meanwhile the original artist gets nothing.

the position of advocating against all copyright protections at all only makes sense for people who are already wealthy enough that they don't need proceeds from their art to survive.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#228
post #41

Considering prices for single work, this must be multi-billion dollar compensation. Take for example 675k paid for 31 songs. So 20k a song. If we estimate book to be say 10MB that would 8 million works. So I think reasonable compensation is something along 163 billion. Not even 10 years of net income. Which I think is entirely fair punishment.

Meta argues that it's fair use, and that they just downloaded, and never seeded, all the torrents.

No they never seeded the essence of it ALL :;))

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#229
post #193
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

i know of a company that poisoned an entire town! thats terrorism if done by an individual. the company still exists, just paid a settlement and carried on...

I agree with your point, but will split hairs on using the word "terrorism". I think that should be reserved for people that commit atrocities for some political aim. I'm fairly sure the company in question (I assume Union Carbide) did not poison the town to advance a political agenda.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#230
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

> Spotify's music library was also pirated in the early days. I want to know more, please enlighten me (anyone who knows). I read the book "The Spotify Play" and it made it seem like the pirated music was an internal-only thing and not something available to customers. Is that true?

Users would upload their copies of the music and spotify would replay them. This was obvious to early users, even if they were only consumers, because of the pirate-shout-out-overlays that were in a lot of the poorer quality releases.

Another interesting note, in the early days of spotify, the app would saturate your upload bandwidth while using it. Given their close ties to utorrent, I always assumed that's how they were affording the bandwidth as well.

Pretty brilliant way to bootstrap I guess; they didn't have to pay for bandwidth or content until they already had contracts in place

Post reply on HN