Live data from Hacker News

AI-Generated Data Can Poison Future AI Models

scientificamerican.com

51–60 of 87 posts

Re: AI-Generated Data Can Poison Future AI Models

#51

Some perspectives from someone working in the image space. These tests don't feel practical - That is, they seem intended to collapse the model, not demonstrate "in the wild" performance. The assumption is that all content is black or white - AI or not AI - and that you treat all content as equally worth retraining on. It offers no room for assumptions around data augmentation, human-guided quality discrimination, or…

This is exactly right. Model collapse does not exist in practice. In fact, LLMs trained on newer web scrapes have increased capabilities thanks to the generated output in their training data. For example, "base" pretrained models trained on scrapes which include generated outputs can 0-shot instruction follow and score higher on reasoning benchmarks. Intentionally produced synthetic training data takes this a step fu…

Ironically Claude 3 appears to have certain "quirks" arguably caused by the fact that its training data contains synthetic data. In one instance (https://twitter.com/DimitrisPapail/status/176477229891207585...), it kept referring to itself as ChatGPT.

Granted, one could argue that this only happened because the API version of Claude doesn't appear to use a system prompt. If that's the case, then the LLM lacks any identity otherwise defined by the initial system prompt, and thus, kind of makes one up.

Nonetheless, point remains, it's kind of interesting to see that in the years since the launch of ChatGPT we're already seeing a tangible impact on publicly available training data. LLMs "know" what ChatGPT is, and may even claim to be it.

Re: AI-Generated Data Can Poison Future AI Models

#52
I'm not sure how much of a risk this is to LLMs in particular, but I feel like we're already seeing the impact on image AI models.

Even though they're getting better at generating hands that make sense and other fine details, you can generally tell that an image is AI generated because it has a certain "style". Can't help but wonder if this is partly due to generated images contaminating the training data and causing subsequent AI image generators to stylistically converge over time.

Re: AI-Generated Data Can Poison Future AI Models

#53

Some perspectives from someone working in the image space. These tests don't feel practical - That is, they seem intended to collapse the model, not demonstrate "in the wild" performance. The assumption is that all content is black or white - AI or not AI - and that you treat all content as equally worth retraining on. It offers no room for assumptions around data augmentation, human-guided quality discrimination, or…

As someone also working in the imaging space, ai generated data is useful solong as it's used carefully.

Specifically, we're implementing AI culled training sets which contain some generated data that then gets reviewed manually for a few specific things, then pushed into our normal training workflows. This makes for a huge speedup versus 100% manual culling and the metrics don't lie, the models continue to improve steadily.

There may be a point where they're poisoned and will collapse, but I haven't seen it yet.

Re: AI-Generated Data Can Poison Future AI Models

#54

Earlier quoted context omitted.

This is exactly right. Model collapse does not exist in practice. In fact, LLMs trained on newer web scrapes have increased capabilities thanks to the generated output in their training data. For example, "base" pretrained models trained on scrapes which include generated outputs can 0-shot instruction follow and score higher on reasoning benchmarks. Intentionally produced synthetic training data takes this a step fu…

What happens if you train a model on nothing but AI-generated output, recursively? Does it eventually get inbred?

Why would you limit a model to be like a brain in a vat? Instead let the model out so people use it, then use the chat logs to fine-tune. A chat room is a kind of environment, there is a human, maybe some tools. The LLM text will generate feedback and right there is a learning signal.

Even without a human, if a LLM has access to code execution it can practice solving coding tasks with runtime feedback. There are many ways a LLM could obtain useful learning signals. After all, we got all our knowledge from the environment as well, in the end there is no other source for knowledge and skills.

Re: AI-Generated Data Can Poison Future AI Models

#56
post #4

I think it's interesting that human minds generally (though not always!) improve when exposed to the output of other human minds. It seems to be the opposite for current LLMs.

This might be my biases speaking, but I have a hunch that there's still more potential for human generated content to poison our minds, than AI.

Re: AI-Generated Data Can Poison Future AI Models

#57

It shouldn't be a problem if you only train on legally acquired data. You will know the authors name and can contact them if you so wish.

I don't think any of the major players could do that for all their data and they are acquiring it legally.

Re: AI-Generated Data Can Poison Future AI Models

#58

I think AI-generated images are worse for training AI generative models than LLMs, since there are so many now on the internet (see Instagram art related hashtags if you want to see nothing but AI art) compared to the quantity of images downloaded prior to 2021 (for those AI that did that). Text will always be more varied than seeing 10m versions of the same ideas that people make for fun. AI text can also be partial…

That's far from unique to instagram. I loathe Stable Diffiusion and co solely because they've utterly FLOODED every cool art-adjacent website with endless mediocre derivative shit. Like there was always low-effort content of course, but holy fuck, there is SO MUCH MORE now. And some of these people are trying to CHARGE for this uninspired junk!!!

I agree with this despite using SD a lot myself. It's fun to use until you realize the majority of people posting stuff generated with it have almost no creativity, all generating the same things over and over again, mostly without any manual work involved. that uncanny realism style with the generic Stable Diffusion face and one of 5 different poses. The number of people putting any sort of effort into it is way, way lower than the number of users thinking they are making art. It's more of a slot machine in the majority of cases

Re: AI-Generated Data Can Poison Future AI Models

#59

Some perspectives from someone working in the image space. These tests don't feel practical - That is, they seem intended to collapse the model, not demonstrate "in the wild" performance. The assumption is that all content is black or white - AI or not AI - and that you treat all content as equally worth retraining on. It offers no room for assumptions around data augmentation, human-guided quality discrimination, or…

This is exactly right. Model collapse does not exist in practice. In fact, LLMs trained on newer web scrapes have increased capabilities thanks to the generated output in their training data. For example, "base" pretrained models trained on scrapes which include generated outputs can 0-shot instruction follow and score higher on reasoning benchmarks. Intentionally produced synthetic training data takes this a step fu…

Wouldn’t reinforcement learning just weigh any nonsense data very low and then spammy garbage doesn’t really affect the model in the end much ? If the model and human experts can’t tell the difference then it’s probably pretty good AI generated data

Re: AI-Generated Data Can Poison Future AI Models

#60

Earlier quoted context omitted.

This is exactly right. Model collapse does not exist in practice. In fact, LLMs trained on newer web scrapes have increased capabilities thanks to the generated output in their training data. For example, "base" pretrained models trained on scrapes which include generated outputs can 0-shot instruction follow and score higher on reasoning benchmarks. Intentionally produced synthetic training data takes this a step fu…

Ironically Claude 3 appears to have certain "quirks" arguably caused by the fact that its training data contains synthetic data. In one instance ( https://twitter.com/DimitrisPapail/status/176477229891207585... ), it kept referring to itself as ChatGPT. Granted, one could argue that this only happened because the API version of Claude doesn't appear to use a system prompt. If that's the case, then the LLM lacks any i…

that is the meat the article tries to cook. the impacts so far aren’t all that negative.

but time flows like a river, and the more shit that gets into it…

poison does not need to be immediately fatal to be fatal. some take a frighteningly long time to work. by the time you know what’s happening, not only is it too late, you have already suffered too much.

does this sound like anything more than a scary story to tell around campfires? not yet.

Post reply on HN