Earlier quoted context omitted.
This is exactly right. Model collapse does not exist in practice. In fact, LLMs trained on newer web scrapes have increased capabilities thanks to the generated output in their training data. For example, "base" pretrained models trained on scrapes which include generated outputs can 0-shot instruction follow and score higher on reasoning benchmarks. Intentionally produced synthetic training data takes this a step fu…
What happens if you train a model on nothing but AI-generated output, recursively? Does it eventually get inbred?
AI-Generated Data Can Poison Future AI Models
81–87 of 87 posts
Re: AI-Generated Data Can Poison Future AI Models
#82Some perspectives from someone working in the image space. These tests don't feel practical - That is, they seem intended to collapse the model, not demonstrate "in the wild" performance. The assumption is that all content is black or white - AI or not AI - and that you treat all content as equally worth retraining on. It offers no room for assumptions around data augmentation, human-guided quality discrimination, or…
Precisely.
Whether content is AI-generate, ghostwriter-generated, monkey-on-keyboard-generated, etc...presumably it is implictly filtered by value/quality.
Garbage AI outputs won't be as popular as good AI outputs. (And the same is true of human ones!)
Re: AI-Generated Data Can Poison Future AI Models
#83Re: AI-Generated Data Can Poison Future AI Models
#84Earlier quoted context omitted.
On what basis are you saying that? I have a model in my head mapping out the world around me, so I know where things are etc and what I can do with all those things. How is that not a world model? Are you using a very strange definition of "world model"?
> On what basis are you saying that? This is a longstanding critique of GOFAI; see Hubert Dreyfuss and Phil Agre. https://pages.gseis.ucla.edu/faculty/agre/critical.html > I have a model in my head mapping out the world around me, so I know where things are etc and what I can do with all those things. No you don't; a map is not the territory, and is necessarily wrong, which means that if you had such a model and were…
What are you talking about, a world model doesn't need to be perfect, nobody said humans has a perfect world model just that we have a world model.
Yes we update the model and adjust when it is wrong, but we still have a world model we use to plan out actions before we do them. I can very accurately predict all the events that will happen when I cook food, how the water will flow etc, sometimes things go wrong and I adjust and fix, but there is no way I could cook food if I didn't have that world model to plan out actions.
> This is a longstanding critique of GOFAI; see Hubert Dreyfuss and Phil Agre.
They don't even mention world model there, I think you are talking about something completely different. No, the mental model humans have of the world isn't the real world, we know that, it is a model of the world, ie a world model. A model isn't the real thing, that is why we call it a model.
Re: AI-Generated Data Can Poison Future AI Models
#85It shouldn't be a problem if you only train on legally acquired data. You will know the authors name and can contact them if you so wish.
If they are scamming and you contact them, of course they will lie.
So how does this work?
Re: AI-Generated Data Can Poison Future AI Models
#86Human created content is also filled with gibberish and false information and random noise… how is AI generated content worse?
> how is AI generated content worse? This is a crucial question. In human society, a feedback loop of nonsense is usually defeated by practical effects in physical reality and experience. The objective of education, for example, is to transmit knowledge and apply reason to important questions. In manipulated social media , there is no check on the nonsense loop. The technology that we currently call A.I. could be use…
Re: AI-Generated Data Can Poison Future AI Models
#87I think it's interesting that human minds generally (though not always!) improve when exposed to the output of other human minds. It seems to be the opposite for current LLMs.