Live data from Hacker News

AI-Generated Data Can Poison Future AI Models

scientificamerican.com

31–40 of 87 posts

Re: AI-Generated Data Can Poison Future AI Models

#31
post #4

I think it's interesting that human minds generally (though not always!) improve when exposed to the output of other human minds. It seems to be the opposite for current LLMs.

I mean it makes sense that (even impressively functional) statistical approximations would degrade when recursed.

If anything I think this just demonstrates yet again that these aren't actually analogous to what humans think of as "minds", even if they're able to replicate more of the output than makes us comfortable.

Re: AI-Generated Data Can Poison Future AI Models

#32
I also wonder what search engines are going to do about all this. Sounds to me, actually, traditional, non-intelligent search might be on its way out, although of course it'll take time. Future search engines will have to be quite adept at trying to figure out whether the text they index is bullshit or not.

Re: AI-Generated Data Can Poison Future AI Models

#34
post #6

Earlier quoted context omitted.

Maybe it's less about "Human VS Robot" and more about exposure to "Original thoughts VS mass-produced average thoughts". I don't think a human mind would be improving if they're in a echo-chamber with no new information. I think the reason the human mind is improving is because we're exposed to new, original and/or different thoughts, that we hadn't considered or come across before. Meanwhile, a LLM will just regurgi…

> I don't think a human mind would be improving if they're in a echo-chamber with no new information If this were true of humans, we would have never made it this far Humans are very capable of looking around themselves and thinking "I can do better than this", and then trying to come up with ways how LLMs are not

> Humans are very capable of looking around themselves and thinking "I can do better than this"

Doesn't this require at least some perspective of what "better than this" means, which you could only know with at least a bit of outside influence in one way or another?

Re: AI-Generated Data Can Poison Future AI Models

#35

Human created content is also filled with gibberish and false information and random noise… how is AI generated content worse?

It is worse, because it is faster - how many incorrect blog articles can a sigle typical writer publish and post on the internet - maybe 1-2 a day if you are a prolific writer?

How many can an AI agent do? Probably hundreds of thousands a day. To me, that is going to be a huge problem - but don't have a solution in mind either.

And then those 100K bad articles posted per day by one person, are used as training data for the next 100K bad/incorrect articles etc - and the problem explodes geometrically.

Re: AI-Generated Data Can Poison Future AI Models

#36
post #4

I think it's interesting that human minds generally (though not always!) improve when exposed to the output of other human minds. It seems to be the opposite for current LLMs.

Humans exhibit very similar behavior. Prolonged sensory deprivation can drive a single individual insane. Fully isolated/monolithic/connected communities easily become detached from reality and are susceptible to mass psychosis. Etc etc etc. Humans need some minimum amount of external data to keep them in check as well.

Re: AI-Generated Data Can Poison Future AI Models

#37

Some perspectives from someone working in the image space. These tests don't feel practical - That is, they seem intended to collapse the model, not demonstrate "in the wild" performance. The assumption is that all content is black or white - AI or not AI - and that you treat all content as equally worth retraining on. It offers no room for assumptions around data augmentation, human-guided quality discrimination, or…

This is exactly right. Model collapse does not exist in practice. In fact, LLMs trained on newer web scrapes have increased capabilities thanks to the generated output in their training data.

For example, "base" pretrained models trained on scrapes which include generated outputs can 0-shot instruction follow and score higher on reasoning benchmarks.

Intentionally produced synthetic training data takes this a step further. For SoTA LLMs the majority of, or all of, their training data is generated. Phi-2 and Claude 3 for example.

Re: AI-Generated Data Can Poison Future AI Models

#38
I think AI-generated images are worse for training AI generative models than LLMs, since there are so many now on the internet (see Instagram art related hashtags if you want to see nothing but AI art) compared to the quantity of images downloaded prior to 2021 (for those AI that did that). Text will always be more varied than seeing 10m versions of the same ideas that people make for fun. AI text can also be partial (like AI-assisted writing) but the images will all be essentially 100% generated.

Re: AI-Generated Data Can Poison Future AI Models

#39
post #4

I think it's interesting that human minds generally (though not always!) improve when exposed to the output of other human minds. It seems to be the opposite for current LLMs.

Reproductive analogy:

A sequence of AI models trained on each other's output gets mutations, which might help or hurt, but if there's one dominant model at any given time then it's like asexual reproduction with only living descendant in each generation (and all the competing models being failures to reproduce). A photocopy of a photocopy of a photocopy — this seems to me to also be the incorrect model which Intelligent Design proponents seem to mistakenly think is how evolution is supposed to work.

A huge number of competing models that never rise to dominance would be more like plants spreading pollen in the wind.

A huge number of AI there are each smart enough to decide what to include in its training set would be more like animal reproduction. The fittest memes survive.

Memetic mode collapses still happen in individual AI (they still happen in humans, we're not magic), but that manifests as certain AI ceasing to be useful and others replacing them economically.

A few mega-minds is a memetic monoculture, fragile in all the same ways as a biological monoculture.

Re: AI-Generated Data Can Poison Future AI Models

#40
post #34

Earlier quoted context omitted.

> I don't think a human mind would be improving if they're in a echo-chamber with no new information If this were true of humans, we would have never made it this far Humans are very capable of looking around themselves and thinking "I can do better than this", and then trying to come up with ways how LLMs are not

> Humans are very capable of looking around themselves and thinking "I can do better than this" Doesn't this require at least some perspective of what "better than this" means, which you could only know with at least a bit of outside influence in one way or another?

Parsimony, explanatory power, and aesthetics. These are things that could be taught to a computer, and I think we will. We had to evolve them too.
Post reply on HN