Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

751–760 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#751
post #185

Earlier quoted context omitted.

You're conflating different problems. Big corporations are too big, they should just not exist. When you have corporations more powerful than the government of the biggest states, it's a bug, not a feature. The IP laws may need rethinking. Saying that they should disappear because big corporations are above the law doesn't help, though. First kill the big corporations, then think about fair laws. Changing the law now…

Economies of scale generate value

And monopolies do the opposite.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#752
post #696

Earlier quoted context omitted.

Uber definitely improved things. When traveling it’s also so much safer than taxis. My brother was robbed at gunpoint in a taxi. My wife had to jump from more than one moving taxi to escape. My ex girlfriend too. My Swiss friend had his camera and wallet stolen. You can have issues with Uber too, but not as frequently because there’s a digital audit trail, you can report them to the platform and the police. The threa…

Were these acts committed by the drivers or somebody else?

Drivers

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#753
post #746

Earlier quoted context omitted.

Uber definitely improved things. When traveling it’s also so much safer than taxis. My brother was robbed at gunpoint in a taxi. My wife had to jump from more than one moving taxi to escape. My ex girlfriend too. My Swiss friend had his camera and wallet stolen. You can have issues with Uber too, but not as frequently because there’s a digital audit trail, you can report them to the platform and the police. The threa…

It didn't improve things for the drivers

That’s not universally true

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#754
post #329

Earlier quoted context omitted.

I just can't get behind the sentiment that the unethical behavior by big companies means I get to access all the content I want for free.

“There is no ethical consumption under capitalism”

Who even thinks that ethical consumption exists under any system? Any of your consumption denies it to others. Some consumption is a necessity of course. We wouldn't speak of something absurd like "ethical breathing".

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#755
post #571
post #470

Earlier quoted context omitted.

Then they'd already be smaller, so there's no reason to make them smaller. Or am I misunderstanding your question?

Okay, they would be smaller, but you said "big corporations should not be able to exist" and they would already be a big corporation with just search--they started this way. Or, just to follow it through, let's say "WidgetBoss LLC" makes a new Widget that every single human has to have, they become the biggest company ever by making one widget. What will you do to make them smaller? Why? I have a big problem with Goo…

I'm not sure if you're in good faith, but I will assume that you are.

> "Literally every billionaire is evil and exploiting blah blah blah"

Nope. Not every billionaire is evil and exploiting blah blah blah. But nobody deserves to be a billionaire, period.

> let's say "WidgetBoss LLC" makes a new Widget that every single human has to have, they become the biggest company ever by making one widget

Which hasn't happened because, obviously, it is not possible to become the biggest company ever by making something trivial.

It is not possible to promote your product by putting it at the top of the search results if you don't own the search engine.

It is not possible to get statistics about popular products in your webstore, copy them and put them at the top of the search results if you can't own both the webstore and the products.

It is not possible to force everybody to use your email provider in order to use their smartphone if you don't own both the email provider and the smartphone OS.

etc.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#756
post #469
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

> If you plug a laptop into a closet at MIT to download some scientific papers you forfeit your life. I’m opposed to copyright and pro-aaronsw, but the state did not kill him.

Absolutely.

1.8 million people are in United States jails today. It isn't a death sentence, and it is a foreseeable consequence of some ethically-appropriate actions.

Supporting folks spending time in jail is a valuable role in any social movement.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#757
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

Exactly. We need leaders with the political will to apply a "financial death penalty" to companies that engage in this kind of brazen behavior. That means all assets seized, the company dissolved, personal assets of executives seized, executives jailed. People running companies should live in mortal fear of ever doing the things that they routinely do today.

Do people even take civics classes anymore? That isn't how any of this works. Political will doesn't allow arbitrary punishments. You would need legislation at very least and that could face issues with the Eighth Amendment. (Which could not be post-facto of course.)

At least you're not calling for jailing all the shareholders....

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#758

Earlier quoted context omitted.

This is starting to get pretty circular. The AI was trained on copyrighted data, so we can make a hypothesis that it would not exist - or would exist in a diminished state - without the copyright infringement. Now, the AI is being used to flood AI bookstores with cheaply produced books, many of which are bad, but are still competing against human authors.

the problem with how circular the argument is is that the essence of there being an actual problem is being taken for granted it's not clear that detriments actually exist, and the benefits are clear

The benefits are not clear: why should an "author" who doesn't want to bother writing a book of their own get to steal the words of people who aren't lazy slackers?

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#759

Earlier quoted context omitted.

the problem with how circular the argument is is that the essence of there being an actual problem is being taken for granted it's not clear that detriments actually exist, and the benefits are clear

The benefits are not clear: why should an "author" who doesn't want to bother writing a book of their own get to steal the words of people who aren't lazy slackers?

It's as much stealing as piracy is stealing, ie none at all. If you disagree, you and I (along with probably many others in this thread) have a fundamental axiomatic incompatibility that no amount of discussion can resolve.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#760
post #335

Earlier quoted context omitted.

What are we actually worried about happening? Are AI-written books getting published? If they start out-competing humans, is that bad? According to most naysayers, they can't do anything original. Are people asking the AI for books? And then hoping it will spit it out a human-written book word for word?

I think the concern goes to the point of copyright to begin with, which is to incentive people to create things. Will the inclusion of copyrighted works in llm training (further) erode that incentive? Maybe, and I think that's a shame if so. But I also don't really think it's the primary threat to the incentive structure in publishing.

People have created for millennia before the modern institution of copyright, so I'm not sure how that's a cogent argument.
Post reply on HN