Live data from Hacker News

Facebook scraped every Australian adult user's public posts to train AI

abc.net.au

191–200 of 267 posts

Re: Facebook scraped every Australian adult user's public posts to train AI

#191
post #110

"They just trust me...Dumb f**s" - Mark Zuckerberg

So, downvoters here are OK with academic results and early resume building disproportionately affecting life trajectory, yet early displays of poor character aren't relevant in holding someone to account in the same way?

This is not a made up quote though I didn't transcribe it exactly, it was an actual message he sent during the early days of building Facebook when asked how he obtained so much personal contact information from Harvard students, so entirely relevant to this context.

https://www.theatlantic.com/magazine/archive/2024/03/faceboo...

Re: Facebook scraped every Australian adult user's public posts to train AI

#192

[dupe] Actual article: https://www.abc.net.au/news/2024-09-11/facebook-scraping-pho... (More discussion: https://news.ycombinator.com/item?id=41508158 )

Ok, we've changed the URL to that from https://www.theverge.com/2024/9/12/24242789/meta-training-ai.... Thanks!

Submitters: "Please submit the original source. If a post reports on something found on another site, submit the latter." - https://news.ycombinator.com/newsguidelines.html

Re: Facebook scraped every Australian adult user's public posts to train AI

#193
post #192

[dupe] Actual article: https://www.abc.net.au/news/2024-09-11/facebook-scraping-pho... (More discussion: https://news.ycombinator.com/item?id=41508158 )

Ok, we've changed the URL to that from https://www.theverge.com/2024/9/12/24242789/meta-training-ai... . Thanks! Submitters: " Please submit the original source. If a post reports on something found on another site, submit the latter. " - https://news.ycombinator.com/newsguidelines.html

..err two posts with the same url only a few days apart? Merge em!

Re: Facebook scraped every Australian adult user's public posts to train AI

#194
post #189

Earlier quoted context omitted.

> Maybe before surveillance tech existed, there was no expectation of privacy in public places, but now that surveillance tech exists, people naturally expect that high-res video of their every move won't be collected, archived and published even if they are in public. currently, that'd be an unrealistic expectation. I'd agree that it would be nice if that wasn't the case but laws need to catch up with technology. Ri…

The topic of expectations reminds me of this article https://spectrum.ieee.org/online-privacy

I think that "shifting baseline syndrome" is a major issue, but on the privacy side of things people don't seem to really understand where we're at currently, and they seem to be very good at lying to themselves about it.

You can find youtube videos of people outright screaming at photographers in public, insisting that no one has a right to take a picture of them without their permission while the entire time they're also standing under surveillance cameras.

When it's in their face they genuinely seem to care about privacy a lot, but they also know the phone in their pocket is so much more invasive in every way. They've been repeatedly told that they're tracked and recorded everywhere.They sign themselves up for it again and again. As long as they don't see it going on right in front of them in the most obvious way possible I guess they can lie to themselves in a way that they can't when they see a man with a camera, but even though on some level they already know that the street photographer is easily the last thing that should concern them, they still get upset to the point where they're screaming in public. I really don't understand it.

Re: Facebook scraped every Australian adult user's public posts to train AI

#195
post #59

I am very likely in the minority here, but I think AI SHOULD be trained on everything that is in the public sphere. I'd be disappointed if it wasn't trained on everything they had access to. If it is trained on private information, then I would have issue with it.

I don't agree because it creates this dilemma for creators: you need to put your work out there to get traction, but if you put your work out there and anything public is fair game, then it will be sampled by a computer and instantly recreated at scale. This might even happen without the operator knowing whose work is being ripped off. Commercial art producers have always ripped off minor artists. They would do it by…

AI isn't ripping off anyone's work. Certainly if it is, it's doing so to a much lesser extent than commissioning an artist to do a piece in another artists style is.

Re: Facebook scraped every Australian adult user's public posts to train AI

#196
post #59

Earlier quoted context omitted.

I don't agree because it creates this dilemma for creators: you need to put your work out there to get traction, but if you put your work out there and anything public is fair game, then it will be sampled by a computer and instantly recreated at scale. This might even happen without the operator knowing whose work is being ripped off. Commercial art producers have always ripped off minor artists. They would do it by…

Well, just as another perspective... I'm not convinced that the philosophy of copyright is a net positive for society. From a certain perspective, all art is theft, and all creativity builds upon preexisting social influences. That's how genres develop, periods, styles... and yes, blatant ripoffs and copycats too. If the underlying goal is to be able to feed creators, maybe society needs better funding models...? The…

> I'm not convinced that the philosophy of copyright is a net positive for society.

It absolutely is, just not in it's current overpowered form.

> where an employer (or other sponsors) pays the living wage for the creator, but the resulting work is then reusable by all.

Some creators want control over their own narrative, and that's entirely reasonable, at least for a limited time.

> I don't buy the argument that nobody would make things if they weren't copyrightable/paid directly.

That was never the argument as far as I'm aware. There are other concerns, like a creator losing all control of their creation before they had a chance to even finish what they wanted to do/tell.

Re: Facebook scraped every Australian adult user's public posts to train AI

#197

I am very likely in the minority here, but I think AI SHOULD be trained on everything that is in the public sphere. I'd be disappointed if it wasn't trained on everything they had access to. If it is trained on private information, then I would have issue with it.

We need to distinguish between modalities of machine intelligence and proceed to set policy tailored to each specific type. (Arguably) we can further include the variable of private, public control; and the orthogonal matter of private or public service.

Machine intelligence is anthropomorphic in utility, that is it serves as either surrogate or substitute for a human cognitive capability. This permits enumeration of AI utility categories. Broadly we can distinguish between creative, knowledgeable, analytical, judicial, predictive, and directing.

As an example use of this approach, consider the case of the AI trained on all public domain material and optionally having had training access to private matter (think Vatican archives). Such an instance should generally not be afforded creative rights, but we would be remiss to restrict its utility as a knowledge base.

The other parameters noted in terms of dual of public|private can of course have bearing on setting type specific constraints.

Re: Facebook scraped every Australian adult user's public posts to train AI

#198
post #180

Can we talk about how most of us haven't read 80% of everything on the internet and yet we are all still better at many basic things than these AIs? At what point do we admit to ourselves that this isn't a sustainable path forward.

Yup. We sure as heck haven't made a lick of progress in the last 2 years. No new, useful technology to see here folks.

Re: Facebook scraped every Australian adult user's public posts to train AI

#199

I am very likely in the minority here, but I think AI SHOULD be trained on everything that is in the public sphere. I'd be disappointed if it wasn't trained on everything they had access to. If it is trained on private information, then I would have issue with it.

Is there anything in the "public sphere" that is not (a) published to the web and (b) under a license that allows Meta to use it for training "AI".

It seems that "AI" is biased toward (1) only bits, and (2) only bits that are published to the internet.

Re: Facebook scraped every Australian adult user's public posts to train AI

#200
post #59

I am very likely in the minority here, but I think AI SHOULD be trained on everything that is in the public sphere. I'd be disappointed if it wasn't trained on everything they had access to. If it is trained on private information, then I would have issue with it.

I don't agree because it creates this dilemma for creators: you need to put your work out there to get traction, but if you put your work out there and anything public is fair game, then it will be sampled by a computer and instantly recreated at scale. This might even happen without the operator knowing whose work is being ripped off. Commercial art producers have always ripped off minor artists. They would do it by…

Where the puck is about to be is very different from where it is. Generative AI hasn't cracked the creativity problem yet. It can generate new art but it can't develop its own style like a human can (from first principles, humans basically caricature high quality video feed).

There is pretty good reason to believe that this will be a solved problem inside a decade. We're moving towards processing video and the power put behind model training keeps increasing. How much is a style worth when computers can just generate 1,000 of them at high speed? It is going to be cheap enough that legal protection is almost irrelevant; ripping off a style will probably be harder than just creating a new original one.

We can wait a bit to find out where the equilibrium is before worrying about what the law should be.

Post reply on HN