Live data from Hacker News

Facebook scraped every Australian adult user's public posts to train AI

abc.net.au

111–120 of 267 posts

Re: Facebook scraped every Australian adult user's public posts to train AI

#111
post #59

Earlier quoted context omitted.

I don't agree because it creates this dilemma for creators: you need to put your work out there to get traction, but if you put your work out there and anything public is fair game, then it will be sampled by a computer and instantly recreated at scale. This might even happen without the operator knowing whose work is being ripped off. Commercial art producers have always ripped off minor artists. They would do it by…

> but if you put your work out there and anything public is fair game, then it will be sampled by a computer and instantly recreated at scale. That's just how the internet works. Don't put something on the internet if you don't want it to be globally distributed and copied. > I personally know two artists who have sued major companies who ripped off their work for ads, and both won million-plus settlements. Ultimatel…

> That's just how the internet works. Don't put something on the internet if you don't want it to be globally distributed and copied.

This is true for average people. Is it true for the wealthy? Is it true for Disney? Does our law acknowledge this truth and ensure equal justice for all?

Re: Facebook scraped every Australian adult user's public posts to train AI

#112

Earlier quoted context omitted.

I guess it's a gray area? It's more like a forum to me, of the pre-Facebook sort, more like Slashdot than reddit. And there's extremely strong moderation here that does the opposite of Facebook: It optimizes against controversy and vitriol rather than encouraging it. We end up with a bunch of nerds mostly talking shop and sometimes complaining about the job market, but that's still far less ragebaitey than most socia…

> It optimizes against controversy and vitriol rather than encouraging it [...] a bunch of nerds mostly talking shop Something doesn't add up here. Nerds talking shop and controversy and vitriol about the their technical preferences is the same thing. Perhaps what you're saying is that controversy and vitriol is only apparent when you're an "innocent bystander" who doesn't have a passion for the subject? Which HN avo…

OK, but I think there's a pretty big difference between "Next.js is too bloated, you should use HTMX" and "so and so group of people are all _____ and they should all be ________, and oh, your mom sucks".

I don't think – I hope, at least – no one is going to start a shooting war over their framework of choice. You can't say the same about much of the content circulating around social media.

(Edit: You added more to your post after I replied. To your point of "Which HN avoids by usually remaining focused on a fairly narrow set of subjects", that's not just my observation, that's the actual guidelines: https://news.ycombinator.com/newsguidelines.html. We self-select into a narrow slice of nerdtalk or we end up getting downvoted or banned from the site. To me that is the big difference between an interest-based forum that generally stays a functional monoculture vs a general social media site that brings diverse strangers together into shouting wars about whatever the controversy du jour is.)

Re: Facebook scraped every Australian adult user's public posts to train AI

#113

I am very likely in the minority here, but I think AI SHOULD be trained on everything that is in the public sphere. I'd be disappointed if it wasn't trained on everything they had access to. If it is trained on private information, then I would have issue with it.

That sounds compelling when you borrow the marketing term "AI" and position the work as part of a sweeping revolution into some beautiful sci-fi future.

It's less compelling when you see the technology as noisy content generators that will flood the network with spam and devour the livelihood and opportunity to learn for low-market artists and programmers.

In the former perspective, you may look at this is "well, what's the best way we can make this happen?" while the latter sees it more like "So you insist on making this happen. Are you sure there's a suitably responsible way for you to do that?"

Re: Facebook scraped every Australian adult user's public posts to train AI

#114
post #59

Earlier quoted context omitted.

I don't agree because it creates this dilemma for creators: you need to put your work out there to get traction, but if you put your work out there and anything public is fair game, then it will be sampled by a computer and instantly recreated at scale. This might even happen without the operator knowing whose work is being ripped off. Commercial art producers have always ripped off minor artists. They would do it by…

Well, just as another perspective... I'm not convinced that the philosophy of copyright is a net positive for society. From a certain perspective, all art is theft, and all creativity builds upon preexisting social influences. That's how genres develop, periods, styles... and yes, blatant ripoffs and copycats too. If the underlying goal is to be able to feed creators, maybe society needs better funding models...? The…

> I'm not convinced that the philosophy of copyright is a net positive for society.

I'm ok with that. But the philosophy of copyrights is not under debate here. All that is being debated is if it should protect small people from big corporations too.

Re: Facebook scraped every Australian adult user's public posts to train AI

#116

Earlier quoted context omitted.

> but if you put your work out there and anything public is fair game, then it will be sampled by a computer and instantly recreated at scale. That's just how the internet works. Don't put something on the internet if you don't want it to be globally distributed and copied. > I personally know two artists who have sued major companies who ripped off their work for ads, and both won million-plus settlements. Ultimatel…

> That's just how the internet works. Don't put something on the internet if you don't want it to be globally distributed and copied. You could make the same argument about paper. "That's just how photocopiers work! If you don't want your creations to be endlessly duplicated and sold, don't write them down!" Heck, you could make the same argument about leaving the house. "That's just how guns work! Don't go out in pu…

> You could make the same argument about paper. "That's just how photocopiers work! If you don't want your creations to be endlessly duplicated and sold, don't write them down!"

No, the argument would be about photocopies, not paper. "That's just how photocopiers work! Don't put something into a photocopier if you don't want photocopies of it." It isn't possible for anyone to access anything on the internet without making copies of that thing. Copies are literally how the internet works.

Shooting everyone who steps outside isn't how guns work either so that also fails as an analogy.

The internet was specifically designed for the global distribution of copies. If that isn't what you want, don't publish your works there.

> That something is technically possible doesn't make it morally right.

Morality is entirely different from how the internet works, but in practice, I don't see anything immoral about making a copy of something. Morality only becomes an issue when it comes to what someone does with that copy.

Re: Facebook scraped every Australian adult user's public posts to train AI

#117

Earlier quoted context omitted.

> but if you put your work out there and anything public is fair game, then it will be sampled by a computer and instantly recreated at scale. That's just how the internet works. Don't put something on the internet if you don't want it to be globally distributed and copied. > I personally know two artists who have sued major companies who ripped off their work for ads, and both won million-plus settlements. Ultimatel…

> That's just how the internet works. Don't put something on the internet if you don't want it to be globally distributed and copied. You could make the same argument about paper. "That's just how photocopiers work! If you don't want your creations to be endlessly duplicated and sold, don't write them down!" Heck, you could make the same argument about leaving the house. "That's just how guns work! Don't go out in pu…

> You could make the same argument about paper.

Most paper doesn't come with Terms and Conditions that everything you write on it belongs to the paper company. I hate Facebook (with a fiery passion) but people gave them their data in exchange for the groundbreaking and unprecedented ability to make friends with another person (which has never been done before). It sucks, but don't use these "free" systems without understanding the sinister dynamics and incentives behind them.

People make the same arguments about the NSA. "They aren't doing anything bad with the data their collecting about every US citizen." Well, at some point they will. Stop borrowing against future freedom for a tiny bit of convenience today.

Re: Facebook scraped every Australian adult user's public posts to train AI

#118
post #75

Earlier quoted context omitted.

What you get from the self-supervised training of a base model is more like "language fluency plus a web of crystallized-knowledge relationships." But also, ML model training is a bit like the stock market: the noise/stupidity in individual examples points in a bunch of random directions, and so ends up cancelling out; while the signal all points in the same direction, and so ends up captured in the distilled model.…

It's still wrong if the majority is wrong, isn't it? Like propaganda on major channels which praise the same thing, is consumed by a major part of the population which then assume that it's the right thing and continue the misinformation, which then propagates to the AI model.

Yes; but in the case of a "common misconception" like this, there's always also a nontrivial minority who do know the "right answer"[1] — and so enough examples of that occur in the training data to enable the model to embed the knowledge-web of "right ideas" (as a niche activation), alongside the "wrong idea" (in its default-mode network).

The "initial fine-tuning of a 'raw' base model to produce a 'generalized pre-trained' base model" process, is commonly talked about in terms of "alignment" — making the model ethical, making it not swear at you, making it refuse to engage with certain content, etc. And really, that part is all optional, with there existing "non-aligned" or "orthogonalized" models that don't have these steps performed on them or have had them reversed, but which are still useful.

But a large part of this initial fine-tuning process, consists of debiasing the model's default activation, moving it away from making associations with "common misconceptions" and toward making associations with "right answers." And this process is crucial to a model being able to reason intelligently — as these common misconceptions aren't coherent in a chain of reasoning they appear in, and so lead to the chain of reasoning falling apart / being non-productive. This is, in large part, the "secret sauce" that makes a model of a given size "more intelligent" than another model of the same size.

Every base model that anyone actually cares about or uses — "aligned" or not — has had some process of de-biasing like this applied to it; or, at least, has short-cutted this process by training on a training dataset generated by or filtered by a model that has already had this de-biasing applied to it, such that the derived training dataset doesn't contain the "common misconceptions" in the first place.

And when OpenAI and Meta brag about using RLHF, a large part of what they mean, is crowdsourcing recognition of long-tail "common misconceptions" at scale, to allow a much more thorough version of this de-biasing process[2].

---

[1] Of course, if nobody in the training data ever demonstrates the "right answer" knowledge/associations, then the model will never learn that knowledge/associations. But then, given that these training datasets usually represent decent samples of the population, a "right answer" being missing entirely would likely mean that nobody on Earth knows the "right answer" — so we humans wouldn't be able to recognize the model was wrong. The stock market can be wrong too, for the same reason.

[2] Which, perhaps surprisingly, can add up to more than the sum of its parts. The more of this human-labelled "common misconception"-response RLHF data you have, the more you can derive patterns from this bias data. You can distil out negative examples that you can use to prompt a model to filter the training dataset; but more interestingly, you can distil positive examples of the sort of structured chains of reasoning that work best to inherently avoid triggering the bias. Where, if you can overlap the activation-space hyperspheres of many such inversion-of-bias examples, then you get, essentially, the hypersphere within the model's activation space that contains its instrumental rationality. You can then just bias the model toward living in that part of activation space as much as possible — and this shoots its apparent reasoning capacity way up.

Re: Facebook scraped every Australian adult user's public posts to train AI

#119

Why is this surprising? They’ve always done this. In fact I’d be surprised if they didn’t do this. Fwiw - llama is free to use. So I guess it’s a good enough return. I don’t use Facebook. I’m not sure if they can peek into WhatsApp messages.

> Why is this surprising? They’ve always done this. In fact I’d be surprised if they didn’t do this. This is such an unconstructive attitude. This is the first time they have publicly admitted it.

Yeah, to me it's part of what press critic Jay Rosen calls The Church of the Savvy. It's kind of a performative cynicism where one tries to gain status by appearing so smart that you're above it all. One can do it with pretty much anything, which ironically means it demonstrates very little actual smarts.

To me it's related to the sort of person you'll see who on every startup failure says how they knew it wasn't going to work. Which again doesn't require particular smarts; most startups fail, so predicting failure doesn't take a genius. What I think is much more interesting is spotting a problem before it's obvious and naming it in advance. Or better, fixing it early on. That doesn't get you many internet points, though.

Re: Facebook scraped every Australian adult user's public posts to train AI

#120

I am very likely in the minority here, but I think AI SHOULD be trained on everything that is in the public sphere. I'd be disappointed if it wasn't trained on everything they had access to. If it is trained on private information, then I would have issue with it.

[deleted]
Post reply on HN