Live data from Hacker News

Facebook scraped every Australian adult user's public posts to train AI

abc.net.au

141–150 of 267 posts

Re: Facebook scraped every Australian adult user's public posts to train AI

#141

I am very likely in the minority here, but I think AI SHOULD be trained on everything that is in the public sphere. I'd be disappointed if it wasn't trained on everything they had access to. If it is trained on private information, then I would have issue with it.

When I speak to my friends, it's a conversation not wholly private - after all, I've shared whatever I'm saying with them - but it certainly isn't wholly public.

In all our conversations, we have and we understand there are degrees of privacy; that which we share with family, that with friends, that with strangers.

When I post on-line, I both expect and expected that my conversations would be between me and the group of people I conversed with. I knew who was reading, and I was fine to write whatever I was writing to that group.

I may be wrong, but I think this is generally how people feel, how they act, what they expect, how they are, as humans. We think about who we are writing to. It does not come naturally to imagine that third parties are listening in, or will listen in, in the decades to come.

This brings us to now, with a third party, reaching back over ten or fifteen years, for absolutely everyone, everywhere, taking copies of everything it can get access to, for its own use, whatever that may be.

I profoundly reject Microsoft, and Google, and all entities and companies which act in such ways, these smiling evils, with their friendly icons and bright colours, happy faces and hundred page T&Cs to utterly obscure and obliterate the truth of their actions.

Re: Facebook scraped every Australian adult user's public posts to train AI

#142

Earlier quoted context omitted.

Their algorithm sure does a good job at filtering those normal people out...

Yes, it does. But that doesn't invalidate the GP.

It just makes it less likely that these normal people will be able to see the other normal posts from other normal people. Early Facebook's Wall was like that, mostly just friends chatting with each other about cats and babies, but then the company purposely started to optimize the timeline for controversy instead and it all went downhill.

It's not that there aren't normal people on there, it's that organic feel-good posts have a harder time gaining traction vs the sea of flamebait and sponsored ads and astroturfed spam. The signal to noise ratio was very, very low by the time I left. I don't know how it is these days...

Re: Facebook scraped every Australian adult user's public posts to train AI

#144
post #65
post #47

Earlier quoted context omitted.

Right, and in exchange users received rather a lot of services for free... photo and video storage etc as just one example, Llama is free to use, etc etc. While I've no sympathy for mishandling private user data as Meta has of course been guilty of in the past, I think users getting 16 years of free service in exchange for their public posts being used in this fashion is not that bad of a deal.

You get photo/video storage worth a fraction of a cent. Llama doesn't have a free to use service. People got 'free service' in exchange for putting ads around content not for this. It's a terrible deal no one asked the users if they want to agree to. It's visiting a website and saving an image and claiming the website owes you for storage.

> Llama doesn't have a free to use service.

Llama is embedded for free in a ton of Meta products right now:

https://about.fb.com/news/2024/04/meta-ai-assistant-built-wi...

> "A better assistant: Thanks to our latest advances with Meta Llama 3, we believe Meta AI is now the most intelligent AI assistant you can use for free"

Want to just use it in a browser?

https://www.meta.ai/

Re: Facebook scraped every Australian adult user's public posts to train AI

#145

Earlier quoted context omitted.

OK, but I think there's a pretty big difference between "Next.js is too bloated, you should use HTMX" and "so and so group of people are all _____ and they should all be ________, and oh, your mom sucks". I don't think – I hope , at least – no one is going to start a shooting war over their framework of choice. You can't say the same about much of the content circulating around social media. (Edit: You added more to…

> there's a pretty big difference between "Next.js is too bloated, you should use HTMX" and "so and so group of people are all _____ and they should all be ________, and oh, your mom sucks". Is there? Perhaps the trouble here is that your examples are too far apart to recognize how they compare? What if Facebook, instead, said "Fat workers are too bloated, you should hire skinny workers"? Or if HN said "so and so pro…

I think it's just the different degree of emotional attachment people to have to these topics. Yeah, people can get a bit worked up about how annoying JS development has become, but not quite to the same level as the major headlines of the day about the latest Middle East controversy/Soviet threat/identity politics thing.

Oh, and my vacuum does suck just fine, thank you very much.

Re: Facebook scraped every Australian adult user's public posts to train AI

#146
post #131

Earlier quoted context omitted.

> but if you put your work out there and anything public is fair game, then it will be sampled by a computer and instantly recreated at scale. That's just how the internet works. Don't put something on the internet if you don't want it to be globally distributed and copied. > I personally know two artists who have sued major companies who ripped off their work for ads, and both won million-plus settlements. Ultimatel…

> That's just how the internet works. Don't put something on the internet if you don't want it to be globally distributed and copied Until now, this has been an acceptable tradeoff because there's some friction to theft. Directly cloning the work is easy, but that also means an artist can sue or DMCA. It also means the original artist's work can go more viral, which, despite the short-term downsides, can help their p…

> With an LLM, it takes milliseconds, and that model will be able to churning out the likes of your work millions of times per day, forever.

AI does cause a lot of problems in terms of scale. The good news is that if AI churns out millions of copies of your copyrighted works you're entitled to compensation for each and every copy. In addition to pushing out copies of copyrighted material, AI is also capable of writing up DMCA notices and legal paperwork.

> With the exception of an LLM directly plagiarizing, the only way to prove it didn't is by not allowing it to train on something. LLMs copy everything and nothing at the same time.

An AI's output should be held to the exact same standard as anyone else's output. If it's close enough to someone else's copyrighted work to be considered infringing then the company using that AI should be liable for copyright infringement the same way they would be if AI had never been involved. AI's ability to produce a large number of infringing works very quickly might even be what causes companies to be more careful about how they use it. Breaking the law at speeds approaching the speed of light isn't a good business model.

Re: Facebook scraped every Australian adult user's public posts to train AI

#147
post #59

I am very likely in the minority here, but I think AI SHOULD be trained on everything that is in the public sphere. I'd be disappointed if it wasn't trained on everything they had access to. If it is trained on private information, then I would have issue with it.

I don't agree because it creates this dilemma for creators: you need to put your work out there to get traction, but if you put your work out there and anything public is fair game, then it will be sampled by a computer and instantly recreated at scale. This might even happen without the operator knowing whose work is being ripped off. Commercial art producers have always ripped off minor artists. They would do it by…

This benefits actually everyone.

If our combined creative work until this point is what turns out to be necessary to kick-start a great shot at abundance (and if you do not believe that, if it's all for nothing, why care at all about the money wasted on models?) it might simply be our societal moral obligation to endorse it -- just as is will be the model creators moral obligation to uphold their end of this deal.

Interestingly, Andrej Karpathy recently described the data we are debating as more or less undesirable to build a better LLM and accidentally good enough to have made it work so far (https://youtu.be/hM_h0UA7upI?t=1045). We'll see about that.

Re: Facebook scraped every Australian adult user's public posts to train AI

#148
post #6

I don't understand how this surprises anyone. You choose to give them your data. It's not free. If you don't want them to have your data, don't give it away.

Easily said in 2024, but 17 years ago? I don't think this was quite so obvious (even amongst technical people).

Re: Facebook scraped every Australian adult user's public posts to train AI

#149

It's OK. Meta is training their AI on hundreds of thousands of posts with photos of veterans with toilet plunger legs celebrating their birthdays in the middle of the street while sitting as sturdy as the Lincoln memorial. The AI brain rot has already begun in this model.

I am truly impressed by how quickly AI generated content has filled up every public space. We've gone from "AI is the future!" to a digital Kessler's syndrome in a short few years.

Not only filled up every public space, but the quality of it all is so crude. Like a bad Pixar animation. It's not like Pixel's "Add the photographer back into photo".

Re: Facebook scraped every Australian adult user's public posts to train AI

#150

Earlier quoted context omitted.

> there's a pretty big difference between "Next.js is too bloated, you should use HTMX" and "so and so group of people are all _____ and they should all be ________, and oh, your mom sucks". Is there? Perhaps the trouble here is that your examples are too far apart to recognize how they compare? What if Facebook, instead, said "Fat workers are too bloated, you should hire skinny workers"? Or if HN said "so and so pro…

I think it's just the different degree of emotional attachment people to have to these topics. Yeah, people can get a bit worked up about how annoying JS development has become, but not quite to the same level as the major headlines of the day about the latest Middle East controversy/Soviet threat/identity politics thing. Oh, and my vacuum does suck just fine, thank you very much.

> not quite to the same level as the major headlines of the day about the latest Middle East controversy/Soviet threat/identity politics thing.

Why do you say that? I share in the tech-minded proclivity towards not having much interest in people, so I admit to being largely out of the loop, but what is there to be worked up about where you wouldn't equally get worked up around some tech-based topic?

It just sounds boring an uninteresting to me. When I have encountered discussions about those topics, all I have is some laughter at how silly the people sound. Just like I'm sure how the "Next.js is awful" conversations sound to the metaphorical grandma.

Post reply on HN