Live data from Hacker News

Brazil data regulator bans Meta from mining data to train AI models

apnews.com

21–30 of 55 posts

Re: Brazil data regulator bans Meta from mining data to train AI models

#21
post #8

This proposal gets made pretty frequently in one form or another, and (at least on HN) seems to usually get struck down on this or that procedural ground. But as the various regulatory and judicial and legislative processes grind through different parts of the modern intellectual property issue made so abundantly legible by the modern AI training data gold rush it seems ever more clear that one way or another, we’re…

> But as the various regulatory and judicial and legislative processes grind through different parts of the modern intellectual property issue made so abundantly legible by the modern AI training data gold rush it seems ever more clear that one way or another, we’re going to get a new social contract on IP. what do you mean by that? as far i'm aware ANYTHING that you publish despite being on the internet or not, if t…

Yes, that's correct. Under the Berne Convention all copyright for a work and any derivatives is held with the author, unless the explicitly disclaim it or another legal provision applies (eg fair use for teaching or parody).

However, does an LLM count as a derivative work or a transformative one? That's something for the lawyers to answer.

Re: Brazil data regulator bans Meta from mining data to train AI models

#23
post #8

Earlier quoted context omitted.

> But as the various regulatory and judicial and legislative processes grind through different parts of the modern intellectual property issue made so abundantly legible by the modern AI training data gold rush it seems ever more clear that one way or another, we’re going to get a new social contract on IP. what do you mean by that? as far i'm aware ANYTHING that you publish despite being on the internet or not, if t…

Yes, that's correct. Under the Berne Convention all copyright for a work and any derivatives is held with the author, unless the explicitly disclaim it or another legal provision applies (eg fair use for teaching or parody). However, does an LLM count as a derivative work or a transformative one? That's something for the lawyers to answer.

what are the opinions on [0]? what's the scene for language rather than image?

also what are the opinions on turning generative AI (that doesn't ask permission to creators) public domain? donation money that surpasses the cost of hosting the work to people/groups "creating" with AI, should be a violation of the license? are you allowed to play with the models in hardware made by for-profit entities, like Nvidia?

[0] https://arxiv.org/abs/2212.03860

Re: Brazil data regulator bans Meta from mining data to train AI models

#24

> Compliance must be demonstrated by the company within five working days from the notification of the decision, and the agency established a daily fine of 50,000 reais ($8,820) for failure to do so. $8,820 * 365 = $3.2 million a year is pretty cheap for Meta to be able to do whatever they want with all the data from all 200 million Brazilians. Their annual net income is $39.10 billion, so 0.008%.

It’s not a Blockbuster Video. They’ll eventually increase the penalty for noncompliance or escalate the kind of punishment.

Are there any cases of this actually happening?

Re: Brazil data regulator bans Meta from mining data to train AI models

#26

This proposal gets made pretty frequently in one form or another, and (at least on HN) seems to usually get struck down on this or that procedural ground. But as the various regulatory and judicial and legislative processes grind through different parts of the modern intellectual property issue made so abundantly legible by the modern AI training data gold rush it seems ever more clear that one way or another, we’re…

> I speak only for myself but plenty of people seem to agree: I don’t mind big companies training on generally available data, I mind the IP-laundering.

This removes the big scary emotional part of the debate. Without this, it's weakened quite a bit.

Re: Brazil data regulator bans Meta from mining data to train AI models

#27
Meta updated its Privacy Policy on June 26, to include in its rights the use of data collected in "Meta Products" for the training of GenAI models. This goes against the interpretation of the local data protection law and as such this note was emitted.

The policy update seems to be global: https://www.facebook.com/privacy/policy

Re: Brazil data regulator bans Meta from mining data to train AI models

#28

Well, joke's on you Brazil. I doubt Meta can mine any more data than they already have. They already have all the names, pictures, face biometrics, social graph, location information, political affiliation, relationships and everything that goes into an advertising profile. What else is needed? They can just use the data already mined, which is probably 99% of everything they will ever need for many years to come. Th…

How do you prove that an AI=>LLM has been trained on specific data? How is this enforceable? How is this anything more than some local politicians looking for some headlines?

Re: Brazil data regulator bans Meta from mining data to train AI models

#29

> Compliance must be demonstrated by the company within five working days from the notification of the decision, and the agency established a daily fine of 50,000 reais ($8,820) for failure to do so. $8,820 * 365 = $3.2 million a year is pretty cheap for Meta to be able to do whatever they want with all the data from all 200 million Brazilians. Their annual net income is $39.10 billion, so 0.008%.

[deleted]

Re: Brazil data regulator bans Meta from mining data to train AI models

#30
Information wants to be free. The ethos of the open web - the levy hacker ethos, has always been about unrestricted access and fair use. When content is published openly online, it inherently invites broad consumption, reproduction, and creative reuse by the public. This principle is deeply rooted in the fair use doctrine as applied to the digital realm.

Fair use is evaluated based on the purpose of use, the nature of the copyrighted work, the amount used, and the effect on the market. These factors generally favor the free use of openly published web content. The transformative nature of many reuses, the public availability of original works, the necessity of using entire works in some cases, and the absence of a traditional market for such content all support this interpretation.

This longstanding practice has driven unprecedented innovation and information dissemination, establishing a social contract between content creators and users that treats open web content as "freeware." Any move to impose strict copyright limitations now would stifle innovation and contradict decades of established legal precedent and digital norms.

Post reply on HN