Live data from Hacker News

AI-Generated Data Can Poison Future AI Models

scientificamerican.com

81–87 of 87 posts

Re: AI-Generated Data Can Poison Future AI Models

#81

Earlier quoted context omitted.

This is exactly right. Model collapse does not exist in practice. In fact, LLMs trained on newer web scrapes have increased capabilities thanks to the generated output in their training data. For example, "base" pretrained models trained on scrapes which include generated outputs can 0-shot instruction follow and score higher on reasoning benchmarks. Intentionally produced synthetic training data takes this a step fu…

What happens if you train a model on nothing but AI-generated output, recursively? Does it eventually get inbred?

Does AlphaZero get inbred?

Re: AI-Generated Data Can Poison Future AI Models

#82

Some perspectives from someone working in the image space. These tests don't feel practical - That is, they seem intended to collapse the model, not demonstrate "in the wild" performance. The assumption is that all content is black or white - AI or not AI - and that you treat all content as equally worth retraining on. It offers no room for assumptions around data augmentation, human-guided quality discrimination, or…

> human-guided quality discrimination

Precisely.

Whether content is AI-generate, ghostwriter-generated, monkey-on-keyboard-generated, etc...presumably it is implictly filtered by value/quality.

Garbage AI outputs won't be as popular as good AI outputs. (And the same is true of human ones!)

Re: AI-Generated Data Can Poison Future AI Models

#84
post #79

Earlier quoted context omitted.

On what basis are you saying that? I have a model in my head mapping out the world around me, so I know where things are etc and what I can do with all those things. How is that not a world model? Are you using a very strange definition of "world model"?

> On what basis are you saying that? This is a longstanding critique of GOFAI; see Hubert Dreyfuss and Phil Agre. https://pages.gseis.ucla.edu/faculty/agre/critical.html > I have a model in my head mapping out the world around me, so I know where things are etc and what I can do with all those things. No you don't; a map is not the territory, and is necessarily wrong, which means that if you had such a model and were…

> No you don't; a map is not the territory, and is necessarily wrong, which means that if you had such a model and were actually relying on it you wouldn't be able to do things you obviously can do in real life.

What are you talking about, a world model doesn't need to be perfect, nobody said humans has a perfect world model just that we have a world model.

Yes we update the model and adjust when it is wrong, but we still have a world model we use to plan out actions before we do them. I can very accurately predict all the events that will happen when I cook food, how the water will flow etc, sometimes things go wrong and I adjust and fix, but there is no way I could cook food if I didn't have that world model to plan out actions.

> This is a longstanding critique of GOFAI; see Hubert Dreyfuss and Phil Agre.

They don't even mention world model there, I think you are talking about something completely different. No, the mental model humans have of the world isn't the real world, we know that, it is a model of the world, ie a world model. A model isn't the real thing, that is why we call it a model.

Re: AI-Generated Data Can Poison Future AI Models

#85

It shouldn't be a problem if you only train on legally acquired data. You will know the authors name and can contact them if you so wish.

What? How do you know the data your buying isn't AI generated by the sellers?

If they are scamming and you contact them, of course they will lie.

So how does this work?

Re: AI-Generated Data Can Poison Future AI Models

#86

Human created content is also filled with gibberish and false information and random noise… how is AI generated content worse?

> how is AI generated content worse? This is a crucial question. In human society, a feedback loop of nonsense is usually defeated by practical effects in physical reality and experience. The objective of education, for example, is to transmit knowledge and apply reason to important questions. In manipulated social media , there is no check on the nonsense loop. The technology that we currently call A.I. could be use…

In addition to different decisions made by individuals, they also can't power a feedback loop 24/7 a kerjillion times per minute.

Re: AI-Generated Data Can Poison Future AI Models

#87
post #4

I think it's interesting that human minds generally (though not always!) improve when exposed to the output of other human minds. It seems to be the opposite for current LLMs.

It's almost as if LLMs and human minds operate entirely differently from each other.
Post reply on HN