Live data from Hacker News

Brazil data regulator bans Meta from mining data to train AI models

apnews.com

41–50 of 55 posts

Re: Brazil data regulator bans Meta from mining data to train AI models

#41

This proposal gets made pretty frequently in one form or another, and (at least on HN) seems to usually get struck down on this or that procedural ground. But as the various regulatory and judicial and legislative processes grind through different parts of the modern intellectual property issue made so abundantly legible by the modern AI training data gold rush it seems ever more clear that one way or another, we’re…

> I don’t mind big companies training on generally available data, I mind the IP-laundering large platforms saw AI and instantly closed their platforms making it hard or impossible for external actors to mine that "generally available data," hurting their own users and the open web in the process, and then they mined the data themselves.

As long as a "user" can access those platforms then that data can and will be mined. The people working on such solutions just dont publish them publicly until the debate is settled. If I can view information on any website, authenticated or not, then I can build a bot that will do the same. This does not mean its being done for nefarious reasons. Simply automating the process of bringing what I consider valuable information to me is enough motivation to do it. In my case, the only profit I make is saving time not manually clicking around to access the data I read over morning coffee.

The internet routes around censorship. Its impossible to hide information as long as its meant to be accessed by a human. If companies want to spend engineering hours putting locks then thats their waste.

Many businesses will fail by wasting time and money creating locks that can and will be circumvented.

I agree that a new social contract is inevitable because the only way to prevent data from being mined is to not produce it to begin with. Period. This I know.

Re: Brazil data regulator bans Meta from mining data to train AI models

#42

Well, joke's on you Brazil. I doubt Meta can mine any more data than they already have. They already have all the names, pictures, face biometrics, social graph, location information, political affiliation, relationships and everything that goes into an advertising profile. What else is needed? They can just use the data already mined, which is probably 99% of everything they will ever need for many years to come. Th…

How do you prove that an AI=>LLM has been trained on specific data? How is this enforceable? How is this anything more than some local politicians looking for some headlines?

That too. They never get audited and if they do, it's in another jurisdiction.

Re: Brazil data regulator bans Meta from mining data to train AI models

#45
post #41

Earlier quoted context omitted.

> I don’t mind big companies training on generally available data, I mind the IP-laundering large platforms saw AI and instantly closed their platforms making it hard or impossible for external actors to mine that "generally available data," hurting their own users and the open web in the process, and then they mined the data themselves.

As long as a "user" can access those platforms then that data can and will be mined. The people working on such solutions just dont publish them publicly until the debate is settled. If I can view information on any website, authenticated or not, then I can build a bot that will do the same. This does not mean its being done for nefarious reasons. Simply automating the process of bringing what I consider valuable inf…

When I was young I used to upload fan art of Naruto to DeviantArt. My badly scanned drawings sucked. Everyone else's sucked. It was cool.

Today DeviantArt has its own AI which it promotes over their own users' work. I've read some threads by artists discussing where to go next, between DA, Instagram, ArtStation, and several other new and likely not much better platforms, and one comment that struck me was someone saying it was just not worth it, and their time was better spent networking offline at a gallery.

AI art might actually kill online art communities.

AI-generated articles might kill online publishing.

AI-generated spam bots might kill social media.

We've taken the Internet for granted as grandma and grandpa joined it. Tomorrow people may just get sick of all these algorithms, let go of their smartphones, and go touch some grass. Then every website is just going to be AI bots regurgitating each others' content ad nauseum.

Humans are on the web because of the reach. If AI-generated content steals all the reach, why would anyone post anything on publicly accessible venues instead of just using private ones?

"A new social contract is inevitable because the only way to prevent data from being mined is to not produce it to begin with." But is it though? You are assuming that "not produce it to begin with" is impossible. I'm afraid it's not impossible and the web experiment is at real danger. Maybe not immediately, but will it survive another 20 years in this environment?

Re: Brazil data regulator bans Meta from mining data to train AI models

#46

This proposal gets made pretty frequently in one form or another, and (at least on HN) seems to usually get struck down on this or that procedural ground. But as the various regulatory and judicial and legislative processes grind through different parts of the modern intellectual property issue made so abundantly legible by the modern AI training data gold rush it seems ever more clear that one way or another, we’re…

LLMs are not a general public benefit. Artists whose work is trained upon by text-to-image models aren't made any more whole just because Meta has to share its weights—it just means it's even cheaper for the folks impersonating them or effortlessly ripping off their style to keep doing so.

Meta really does not need to be subsidized when they have so many resources at hand—if LLMs are really hard to train without that much data, then perhaps that's a flaw with the approach instead of something the world has to accommodate.

Re: Brazil data regulator bans Meta from mining data to train AI models

#47
post #8

Earlier quoted context omitted.

> But as the various regulatory and judicial and legislative processes grind through different parts of the modern intellectual property issue made so abundantly legible by the modern AI training data gold rush it seems ever more clear that one way or another, we’re going to get a new social contract on IP. what do you mean by that? as far i'm aware ANYTHING that you publish despite being on the internet or not, if t…

Yes, that's correct. Under the Berne Convention all copyright for a work and any derivatives is held with the author, unless the explicitly disclaim it or another legal provision applies (eg fair use for teaching or parody). However, does an LLM count as a derivative work or a transformative one? That's something for the lawyers to answer.

This has an easy answer — it’s just not the one that people who desperately want to use LLMs for copyright washing want to hear.

The test for what constitutes a derivative work has not changed; it’s the same whether a single human author produced something, or a team of humans, or an LLM. It will be up to a court to decide whether a work is similar enough to be considered derivative.

If an LLM spits out a verbatim copy, that’s obviously infringement. But if the LLM spits out something similar? Well if the LLM spits out something like George Harrison’s My Sweet Lord [1], a court may well decide that it’s derivative of He’s So Fine. Especially if the LLM “subconsciously” “knew” about He’s So Fine because it was part of the training corpus.

[1] https://en.wikipedia.org/wiki/My_Sweet_Lord#Copyright_infrin...

Re: Brazil data regulator bans Meta from mining data to train AI models

#48

This proposal gets made pretty frequently in one form or another, and (at least on HN) seems to usually get struck down on this or that procedural ground. But as the various regulatory and judicial and legislative processes grind through different parts of the modern intellectual property issue made so abundantly legible by the modern AI training data gold rush it seems ever more clear that one way or another, we’re…

> training on roughly “the commons”,

The proposal in the article, however, is not about "the commons", it's about content that the users themselves produced, and then they voluntarily gave permission to Meta to use.

Or are you saying that if I produce some type of material, I shouldn't be able to license it for someone else to use it freely?

Re: Brazil data regulator bans Meta from mining data to train AI models

#49

> Compliance must be demonstrated by the company within five working days from the notification of the decision, and the agency established a daily fine of 50,000 reais ($8,820) for failure to do so. $8,820 * 365 = $3.2 million a year is pretty cheap for Meta to be able to do whatever they want with all the data from all 200 million Brazilians. Their annual net income is $39.10 billion, so 0.008%.

This is the heads up fine _just_ to update the privacy policy. It's a mere warning.

The fine for each privacy infraction is 2% of the company's last year earnings, limited to 50 million BRL (~9 million USD). If 500 brazillians had their privacy violated by a platform, that platform needs to pay 500 of these fines once per day until it is fixed. There's also all sorts of extra punishments for not fixing it in time (like mandatory suspension of services).

Facebook is not forbidden to use your data for AI. It can do so, as long as it provide means for you to delete it. A button to clean your data, for example. That would be legal. We know for LLMs is not that easy though.

Re: Brazil data regulator bans Meta from mining data to train AI models

#50
post #39

Earlier quoted context omitted.

So your plan is to regulate it into complete commercial unviability, where the only source of funding were government bureaucrats? How often does this strategy pay off?

How commercially viable is a public library?

It only exists because it's commercially viable to print and sell books. Also, how much capital does a single library and a single frontier llm require?
Post reply on HN