Live data from Hacker News

Brazil data regulator bans Meta from mining data to train AI models

apnews.com

1–10 of 55 posts

Re: Brazil data regulator bans Meta from mining data to train AI models

#3
This proposal gets made pretty frequently in one form or another, and (at least on HN) seems to usually get struck down on this or that procedural ground.

But as the various regulatory and judicial and legislative processes grind through different parts of the modern intellectual property issue made so abundantly legible by the modern AI training data gold rush it seems ever more clear that one way or another, we’re going to get a new social contract on IP.

Leaving aside for a moment the thicket of laws, precedents, jurisdictions, and regulatory inertia: we can vote with our feet as both customers and contributors for common sense now.

So how about the following compromise: promote innovation by liberalizing the posture around training on roughly “the commons”, but insist that the resulting weights are likewise available to the public. Why do I have to take someone’s word for it that they’ve got a result around superposition or whatever on mech interp? I’d like to see it work given it’s everyone’s data pushing those weights.

I speak only for myself but plenty of people seem to agree: I don’t mind big companies training on generally available data, I mind the IP-laundering. Compete on cost, compete on value-added software stacks, compete on vertical integration. There is lots of money to be made building a better mousetrap in terms of code and infrastructure and product innovation.

Conduct the research in the open. None of this would be possible without an ocean of research and data subsidized in whole or in part by the public. Asserting any form of ownership over the result might end up being legal, but it will never be ethical.

Meta isn’t perfect on this stuff, but they’re by far the actor pulling the conversation in that direction. Let’s encourage them to continue pushing the pace on stuff like LLaMA 3.

Re: Brazil data regulator bans Meta from mining data to train AI models

#4
> Compliance must be demonstrated by the company within five working days from the notification of the decision, and the agency established a daily fine of 50,000 reais ($8,820) for failure to do so.

$8,820 * 365 = $3.2 million a year is pretty cheap for Meta to be able to do whatever they want with all the data from all 200 million Brazilians. Their annual net income is $39.10 billion, so 0.008%.

Re: Brazil data regulator bans Meta from mining data to train AI models

#5

> Compliance must be demonstrated by the company within five working days from the notification of the decision, and the agency established a daily fine of 50,000 reais ($8,820) for failure to do so. $8,820 * 365 = $3.2 million a year is pretty cheap for Meta to be able to do whatever they want with all the data from all 200 million Brazilians. Their annual net income is $39.10 billion, so 0.008%.

It’s not a Blockbuster Video. They’ll eventually increase the penalty for noncompliance or escalate the kind of punishment.

Re: Brazil data regulator bans Meta from mining data to train AI models

#6

This proposal gets made pretty frequently in one form or another, and (at least on HN) seems to usually get struck down on this or that procedural ground. But as the various regulatory and judicial and legislative processes grind through different parts of the modern intellectual property issue made so abundantly legible by the modern AI training data gold rush it seems ever more clear that one way or another, we’re…

> So how about the following compromise: promote innovation by liberalizing the posture around training on roughly “the commons”, but insist that the resulting weights are likewise available to the public.

How much would you personally invest in a startup which would spend billions of dollars on a compute cluster only to release the weights publicly after the training is complete?

Re: Brazil data regulator bans Meta from mining data to train AI models

#7
post #6

This proposal gets made pretty frequently in one form or another, and (at least on HN) seems to usually get struck down on this or that procedural ground. But as the various regulatory and judicial and legislative processes grind through different parts of the modern intellectual property issue made so abundantly legible by the modern AI training data gold rush it seems ever more clear that one way or another, we’re…

> So how about the following compromise: promote innovation by liberalizing the posture around training on roughly “the commons”, but insist that the resulting weights are likewise available to the public. How much would you personally invest in a startup which would spend billions of dollars on a compute cluster only to release the weights publicly after the training is complete?

think of the Investors !

Re: Brazil data regulator bans Meta from mining data to train AI models

#8

This proposal gets made pretty frequently in one form or another, and (at least on HN) seems to usually get struck down on this or that procedural ground. But as the various regulatory and judicial and legislative processes grind through different parts of the modern intellectual property issue made so abundantly legible by the modern AI training data gold rush it seems ever more clear that one way or another, we’re…

> But as the various regulatory and judicial and legislative processes grind through different parts of the modern intellectual property issue made so abundantly legible by the modern AI training data gold rush it seems ever more clear that one way or another, we’re going to get a new social contract on IP.

what do you mean by that? as far i'm aware ANYTHING that you publish despite being on the internet or not, if there isn't a copyright notice, you should assume -> "all rights reserved"

Re: Brazil data regulator bans Meta from mining data to train AI models

#9

This proposal gets made pretty frequently in one form or another, and (at least on HN) seems to usually get struck down on this or that procedural ground. But as the various regulatory and judicial and legislative processes grind through different parts of the modern intellectual property issue made so abundantly legible by the modern AI training data gold rush it seems ever more clear that one way or another, we’re…

Reforming the social contract around IP became infinitely difficult the moment normal people started calling it "intellectual property", forming a tangled mess of legal, moral, and ethical ideas in most people's minds.

Your compromise is exactly the situation I desire but seems untenable to most people.

Re: Brazil data regulator bans Meta from mining data to train AI models

#10
post #6

This proposal gets made pretty frequently in one form or another, and (at least on HN) seems to usually get struck down on this or that procedural ground. But as the various regulatory and judicial and legislative processes grind through different parts of the modern intellectual property issue made so abundantly legible by the modern AI training data gold rush it seems ever more clear that one way or another, we’re…

> So how about the following compromise: promote innovation by liberalizing the posture around training on roughly “the commons”, but insist that the resulting weights are likewise available to the public. How much would you personally invest in a startup which would spend billions of dollars on a compute cluster only to release the weights publicly after the training is complete?

I'm gladly investing portion of my wages to such efforts, alongside e.g. education, healthcare, childcare and infrastructure. And I don't even want any monetary ROI from my investment!
Post reply on HN