Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

851–860 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#851

Earlier quoted context omitted.

Exactly. We need leaders with the political will to apply a "financial death penalty" to companies that engage in this kind of brazen behavior. That means all assets seized, the company dissolved, personal assets of executives seized, executives jailed. People running companies should live in mortal fear of ever doing the things that they routinely do today.

Do people even take civics classes anymore? That isn't how any of this works. Political will doesn't allow arbitrary punishments. You would need legislation at very least and that could face issues with the Eighth Amendment. (Which could not be post-facto of course.) At least you're not calling for jailing all the shareholders....

Political will can be used to pass legislation and ultimately to change the Eighth Amendment.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#852

Earlier quoted context omitted.

> Could make interesting case law. Yeah, to perpetuate this system where only those who can afford lawyers get to benefit

Since it’s case law, everyone would benefit from the precedent

The last time the US Supreme Court decided on copyright law, they basically said "we like what Google is doing with Android, so what they did is fine".

So, no, not necessarily.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#853

Is there a concept in the legal system of first-come-first-served that could be used as precedent? What I mean is: when someone is prosecuted for copyright infringement, but Meta isn't, then could the case be put on hold until Meta is found guilty and pays a fine? Also maybe the fine on the later case would have to be proportional to the prior case. So if Meta pays $1 per infringement, the penalty might be $1 for tor…

Lawyers (and hence, judges) are really good at arguing why the earlier case does not apply in a present case, even if most reasonable people would think the two cases are essentially the same.

It's a fundamental part of lawyer training, and if they want to let BigCorp go and bring the hammer down on the little guy, they can make up a hundred reasons for it.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#854

For some misterious reason I can't see Zuckerberg in front of a judge facing 50 years imprisonment. Anyone can? I truly hope that whoever takes the case goes after Meta with 1000 times the pressure that was put on Swartz, but honestly I don't expect much just as the top comment precisly expressed. And if we are going to be fair please also let's not forget about the other usual suspects, or anyone thinks they are fal…

There are other countries than the US though and if rightsholders wish to sue, lawsuits can happen there too. Several EU countries, Switzerland, South Korea, Japan, etc. are viable countries to sue from. Even in Japan which has a law specifically permitting training on copyrighted material you must still obtain it legally-- i.e. you must license it.

That's irrelevant. Switzerland (for example) isn't going to arrest Zuckerberg and put him in jail for this either.

Nobody will.

But if you're operating a site called Pirate Bay or something like that and it's not earning billions of dollars, expect countries to chase you across the globe trying to arrest you.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#855

Earlier quoted context omitted.

The reason there was no political will to punish Airbnb and Uber for violating the law was that initially they were subsidized with VC money and so were able to undercut traditional hotels and taxis on price. In the world of tradable goods, pricing below cost with the intent of putting competition out of business so you can raise prices later is known as "dumping" and is itself illegal.

I rooted for Uber to smash the Taxi cartels. Let us not forget that Taxi Cartels were also insidious beasts. Taxi drivers abused their walled garden with their price gouging by taking longer routes, refusal to take a credit card, and extremely poorly maintained fleets of vehicles. I have had mostly good experiences with Uber, whereas I had experiences that mostly bordered on general condescension toward me whenever I…

You should look at what’s going on in Cancun Mexico with Uber and the taxis right now. Your description in the first sentence is quite accurate.

Your comment also points out the power that regulation has to enforce and protect monopolies. I’m not saying all regulation is bad obviously, but I think we can see exactly what effect it had on the taxi industry, and I’m sure glad that Uber managed to disrupt it.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#856
post #740

Something tells me uncle Donald will exonerate his new favourite lapdog from any criminal or civil liability.

IANAL but the pardon power (A) only extends to criminal punishments, not civil liabilities and (B) copyright lawsuits can be launched by anybody, not just the Department of Justice. So, barring further Might Makes Right shit--which I'm not willing to fully rule out--Trump can't fully shield Zuckerberg et al.

How unlikely it is for Trump to declare AI national security and simply make it lawless playground fro Zuck & Co.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#857
post #8

Really curious what the judges are going to do here. Horse has functionally bolted on this already I’m guessing slap on wrist despite courts going after individual for a couple of movies torrented pretty hard

Is there any other possible outcome than a fine? That too one which will not really affect Meta's overall earnings

> Is there any other possible outcome than a fine?

Yes, of course.

It's quite possible that judges realize that if they restrict training data to licensed materials, LLMs will become stupid and China will overtake the US to become the leader in AI, and because that can't happen, they'll make up some reason to make training on unlicensed data legal. It's definitely fair use!

I'm not even joking. Last time the US Supreme Court basically said "Android is too important, we have to declare its use of Java API fair use."

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#858
post #515

Earlier quoted context omitted.

> If you plug a laptop into a closet at MIT to download some scientific papers you forfeit your life. In case anybody here doesn't know, that's a reference to Aaron Swartz, an activist (and Reddit co-founder) that was risking 35 years in prison and a $1 million fine just for downloading a lot of academic papers from JSTOR. He eventually took his life because of the pressure. May his soul rest in peace.

Except he was offered 6 months in a plea bargain, which he declined because he wanted a trial. Whether 6 months was reasonable punishment for "plug a laptop into a closet at MIT to download some scientific papers" is another matter, but "you forfeit your life" or "35 years in prison and a $1 million fine " is massively misleading.

Wild to see the concept of the plea bargain being defended. It's a blatant retraction of the due process clause from the bill of rights, and here you are defending it.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#859

Earlier quoted context omitted.

It's as much stealing as piracy is stealing, ie none at all. If you disagree, you and I (along with probably many others in this thread) have a fundamental axiomatic incompatibility that no amount of discussion can resolve.

Stealing is not the right word perhaps, but it is bad, and this should be obvious. Because if you take the limit of these arguments as they approach infinity, it all falls apart. For piracy, take switch games. Okay, pirating Mario isn't stealing. Suppose everyone pirates Mario. Then there's no reason to buy Mario. Then Nintendo files bankruptcy. Then some people go hungry, maybe a few die. Then you don't have a switc…

> Because if you take the limit of these arguments as they approach infinity, it all falls apart.

Not everyone is a Kantian, who has the moral philsophy you are talking about, the categorical imperative. See this [0] for a list of criticisms to said philosophy.

> In a vacuum making an AI book is whatever. In the context of humanity and pushing this to it's limits, we can't even begin to comprehend the consequences. I'm talking crimes against humanity beyond your wildest dreams. If you don't know what I'm talking about, you haven't thought long enough and creatively enough.

Not really a valid argument, again it's circular in reasoning with a lot of empty claims with no actual reasoning, why exactly is it bad? Just saying "you haven't thought long enough and creatively enough" does not cut it in any serious discussion, the burden of substantiating your own claim is on you, not the reader, because (to take your own Kantian argument) anyone you've debating could simply terminate the conversation by accusing you of not thinking about the problem deep enough, meaning that no one actually learns anything at all when everyone is shifting the burden of proof to everyone else.

[0] https://en.wikipedia.org/wiki/Categorical_imperative#Critici...

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#860
post #436

Earlier quoted context omitted.

Nobody has responded to me with anything about how authors are harmed, so I don't really get who we're protecting here. It feels more like we just want to punish people, particularly rich people, particularly if they get away with stuff we're afraid to try.

> Nobody has responded to me with anything about how authors are harmed The same way good law-abiding folk are harmed when Heroin is introduced to the community. Then those people won't be able to lend you a cup of sugar, and may well start causing problems. AI books take off and are easy to digest, and before long your user base is quite literally too stupid to buy and read your book even if they wanted. And, for th…

> know to be morally reprehensible

In your opinion, not to everyone. There has been no actual argument as to why it's supposedly "morally reprehensible."

Post reply on HN