Live data from Hacker News

Brazil data regulator bans Meta from mining data to train AI models

apnews.com

31–40 of 55 posts

Re: Brazil data regulator bans Meta from mining data to train AI models

#31
post #8

Earlier quoted context omitted.

> But as the various regulatory and judicial and legislative processes grind through different parts of the modern intellectual property issue made so abundantly legible by the modern AI training data gold rush it seems ever more clear that one way or another, we’re going to get a new social contract on IP. what do you mean by that? as far i'm aware ANYTHING that you publish despite being on the internet or not, if t…

Yes, that's correct. Under the Berne Convention all copyright for a work and any derivatives is held with the author, unless the explicitly disclaim it or another legal provision applies (eg fair use for teaching or parody). However, does an LLM count as a derivative work or a transformative one? That's something for the lawyers to answer.

obligatory IANAL, but seeing LLMs:

- regurgitate entire passages word for word, until that behavior is publicized and quickly RLHF'd away

- rip github repos almost entirely (some new Sonnet 3.5 demos Anthropic employees were bragging about on Twitter were basically 1:1 to a person's public repo)

It seems clear to me that not only can copyrighted work be retained and returned in near entirety by the architectures that undergird current frontier models, but the engineers working on these models will readily confuse a model regurgitating work to be "creating novel work".

Re: Brazil data regulator bans Meta from mining data to train AI models

#32

> Compliance must be demonstrated by the company within five working days from the notification of the decision, and the agency established a daily fine of 50,000 reais ($8,820) for failure to do so. $8,820 * 365 = $3.2 million a year is pretty cheap for Meta to be able to do whatever they want with all the data from all 200 million Brazilians. Their annual net income is $39.10 billion, so 0.008%.

It’s not a Blockbuster Video. They’ll eventually increase the penalty for noncompliance or escalate the kind of punishment.

with half the fine, Meta buys the entire Brazilian supreme court and suddenly: no fines, no jail and everyone will be happy.

Re: Brazil data regulator bans Meta from mining data to train AI models

#33

This proposal gets made pretty frequently in one form or another, and (at least on HN) seems to usually get struck down on this or that procedural ground. But as the various regulatory and judicial and legislative processes grind through different parts of the modern intellectual property issue made so abundantly legible by the modern AI training data gold rush it seems ever more clear that one way or another, we’re…

> I don’t mind big companies training on generally available data, I mind the IP-laundering

large platforms saw AI and instantly closed their platforms making it hard or impossible for external actors to mine that "generally available data," hurting their own users and the open web in the process, and then they mined the data themselves.

Re: Brazil data regulator bans Meta from mining data to train AI models

#34
post #8

Earlier quoted context omitted.

> But as the various regulatory and judicial and legislative processes grind through different parts of the modern intellectual property issue made so abundantly legible by the modern AI training data gold rush it seems ever more clear that one way or another, we’re going to get a new social contract on IP. what do you mean by that? as far i'm aware ANYTHING that you publish despite being on the internet or not, if t…

Yes, that's correct. Under the Berne Convention all copyright for a work and any derivatives is held with the author, unless the explicitly disclaim it or another legal provision applies (eg fair use for teaching or parody). However, does an LLM count as a derivative work or a transformative one? That's something for the lawyers to answer.

>> derivative work or a transformative one?

This isn't solved even for humans. There are trials that clear misunderstandings about fair use. (Every developer here has heard of this one: https://en.m.wikipedia.org/wiki/Google_LLC_v._Oracle_America....)

Artificial Intelligence currently has no concept of responsibility (not legal, not ethical), and it will never have existential threats derived from law. The only way that I can think of, as of right now, is that every single product touched by AI must have a human who is legally responsible for it.

Re: Brazil data regulator bans Meta from mining data to train AI models

#36

> Compliance must be demonstrated by the company within five working days from the notification of the decision, and the agency established a daily fine of 50,000 reais ($8,820) for failure to do so. $8,820 * 365 = $3.2 million a year is pretty cheap for Meta to be able to do whatever they want with all the data from all 200 million Brazilians. Their annual net income is $39.10 billion, so 0.008%.

It’s not a Blockbuster Video. They’ll eventually increase the penalty for noncompliance or escalate the kind of punishment.

[deleted]

Re: Brazil data regulator bans Meta from mining data to train AI models

#37
post #18

This proposal gets made pretty frequently in one form or another, and (at least on HN) seems to usually get struck down on this or that procedural ground. But as the various regulatory and judicial and legislative processes grind through different parts of the modern intellectual property issue made so abundantly legible by the modern AI training data gold rush it seems ever more clear that one way or another, we’re…

If I post something publically I'm fine with everyone being able to use it for AI and other data mining. But I'm not ok for a single company only to benefit. And definitely not to sell my public data. I'm looking at you reddit.

[deleted]

Re: Brazil data regulator bans Meta from mining data to train AI models

#38

> Compliance must be demonstrated by the company within five working days from the notification of the decision, and the agency established a daily fine of 50,000 reais ($8,820) for failure to do so. $8,820 * 365 = $3.2 million a year is pretty cheap for Meta to be able to do whatever they want with all the data from all 200 million Brazilians. Their annual net income is $39.10 billion, so 0.008%.

I hereby also ban Facebook from accessing my data to train AI. I'm sure it would be a big hassle to carefully exclude me from all their models, but I'm generously offering an alternative non-compliance fine of only $50/day, paid to my bank account.

Re: Brazil data regulator bans Meta from mining data to train AI models

#39

Earlier quoted context omitted.

I'm gladly investing portion of my wages to such efforts, alongside e.g. education, healthcare, childcare and infrastructure. And I don't even want any monetary ROI from my investment!

So your plan is to regulate it into complete commercial unviability, where the only source of funding were government bureaucrats? How often does this strategy pay off?

How commercially viable is a public library?

Re: Brazil data regulator bans Meta from mining data to train AI models

#40
Honestly, I'm rather frustrated by the HN discourse on this topic.

TFA (with emphasis added):

> Brazil’s national *data protection* authority determined on Tuesday that Meta, the parent company of Instagram and Facebook, cannot use data originating in the country to train its artificial intelligence.

> The decision stems from “the imminent risk of serious and irreparable or difficult-to-repair damage to the fundamental rights of the affected data subjects,” the agency said in the nation’s official gazette.

https://www.theregister.com/2024/06/14/meta_eu_privacy/ (with emphasis added):

> The decision to halt AI training using EU content follows complaints to *data protection* agencies in 11 European countries – and those agencies, led by Ireland, telling the Facebook giant to scrap the slurp.

While there is no shortage of IP, licensing, and copyright moral quandaries in training LLMs and their ilk, Meta/FB is not getting regulated on those grounds! They are getting regulated on privacy issues. It's even there on The Register path.

I'm seeing a lot of comments in these threads about IP, copyright, and licensing---which, please do take note, are well-defined legal terms and are not to be used interchangeably---but all that is irrelevant because that is not the question Meta is being made to answer for.

Even more frustrating are threads/arguments to what "irrevocable (copy)rights" you give FB per their TOS without even bothering to cite the relevant bits of the TOS to prove their point. Exercise to the reader: prove/disprove that [a] FB users retain copyright of their content even when posted to FB and [b] you are merely licensing FB to specific (not universal!) uses of your content posted in their platform and [c] said license is revocable any time. The astute reader is referred to the Berne Convention but Facebook's TOS will also do just fine. Standard question, one point per answer.

Bonus point question: if you have proven the points above, what action allows you to revoke the license you have granted FB?

(Of course, end of the day, I'm again playing lawyer in an online forum. I'm no better than anyone else here what do I know.)

Post reply on HN