Live data from Hacker News

Facebook scraped every Australian adult user's public posts to train AI

abc.net.au

91–100 of 267 posts

Re: Facebook scraped every Australian adult user's public posts to train AI

#91

Earlier quoted context omitted.

You don’t consider HN social media?

I guess it's a gray area? It's more like a forum to me, of the pre-Facebook sort, more like Slashdot than reddit. And there's extremely strong moderation here that does the opposite of Facebook: It optimizes against controversy and vitriol rather than encouraging it. We end up with a bunch of nerds mostly talking shop and sometimes complaining about the job market, but that's still far less ragebaitey than most socia…

[deleted]

Re: Facebook scraped every Australian adult user's public posts to train AI

#92

Funny to think that the distillation of 16 years of Facebook posts is now considered "intelligence."

Amusing but a large part of the training data other LLM models is social media platforms. I mean hop on chat GPT right now and ask it to come up with any kind of novel joke. I can almost guarantee it will either be some kind of word play or other equally low hanging kind of pun that you would find on Reddit fifty replies deep. Where's the Mensa member only social media dating platform for a properly erudite and snobb…

Google scholar ?

Re: Facebook scraped every Australian adult user's public posts to train AI

#95

Earlier quoted context omitted.

> but if you put your work out there and anything public is fair game, then it will be sampled by a computer and instantly recreated at scale. That's just how the internet works. Don't put something on the internet if you don't want it to be globally distributed and copied. > I personally know two artists who have sued major companies who ripped off their work for ads, and both won million-plus settlements. Ultimatel…

> That's just how the internet works. Don't put something on the internet if you don't want it to be globally distributed and copied. Or we could be ethical and encourage others to be ethical.

I see you're one of the ones that wouldn't download a car.

Re: Facebook scraped every Australian adult user's public posts to train AI

#96

Earlier quoted context omitted.

I wonder how much safety work they have to do specifically because of this. I’d imagine their model might have a fairly paranoid, slightly racist bias if not. Particularly as younger demographics shifted away from FB in the last decade.

Actually if you want to train a model that can recognize bad things you want to have those bad things in the language model training data, otherwise it won’t see the characteristics of those things and it will later struggle to recognize them in later training stages.

Presence may be necessary, but researcher-driven weighting of different sources of content can still introduce bias. For instance, [0] suggests (sources are unclear) that OpenAI boosted by 5x the weight of their WebText2 dataset, which consists of sites linked to by upvoted Reddit comments. Reddit, in this sense, with all the biases of its various communities, was artificially elevated in importance. (Per [1], there were well-thought-through reasons for this around previous failures due to overreliance on Common Crawl, but it's nonetheless a choice that was made by humans to go in this direction.)

[0] https://gregoreite.com/drilling-down-details-on-the-ai-train...

[1] https://insightcivic.s3.us-east-1.amazonaws.com/language-mod...

Re: Facebook scraped every Australian adult user's public posts to train AI

#97
post #75

Funny to think that the distillation of 16 years of Facebook posts is now considered "intelligence."

What you get from the self-supervised training of a base model is more like "language fluency plus a web of crystallized-knowledge relationships." But also, ML model training is a bit like the stock market: the noise/stupidity in individual examples points in a bunch of random directions, and so ends up cancelling out; while the signal all points in the same direction, and so ends up captured in the distilled model.…

People are frequently wrong in the same way in environments where they are easily influenced by each other like social media. That's part of why these models, especially early on, exhibited so many racial and gender biases.

Re: Facebook scraped every Australian adult user's public posts to train AI

#99
post #59

I am very likely in the minority here, but I think AI SHOULD be trained on everything that is in the public sphere. I'd be disappointed if it wasn't trained on everything they had access to. If it is trained on private information, then I would have issue with it.

I don't agree because it creates this dilemma for creators: you need to put your work out there to get traction, but if you put your work out there and anything public is fair game, then it will be sampled by a computer and instantly recreated at scale. This might even happen without the operator knowing whose work is being ripped off. Commercial art producers have always ripped off minor artists. They would do it by…

> Why would we embrace this now that a computer can do it and there's a level of deniability?

Generally I don't think people are arguing that copyright law should be more lenient to AI than it is to humans. If your work gets ripped off (a substantially similar copy not covered by fair use) you can sue regardless of tools used in its creation.

Question would be whether machine learning, unlike human learning, should be treated as copyright infringement. There are differences and the law does not inherently need to treat them the same, but it could.

As to why it should: I think there's huge benefit across a large range of industries to web-scale pretraining and foundation models, and I'd like it to remain accessible to open-source groups or smaller companies without huge data moats. Realistically I think the alternative would likely just benefit Getty/Universal with near-identical outcomes for most actual artists.

When the very basis of copyright is for the "progress of sciences and useful arts", it seems backwards to use it in a way that would set back advances in language translation, malware/spam/DDoS filtering, defect detection, voice dictation/transcription, medical image segmentation, etc.

Post reply on HN