Live data from Hacker News

AI-Generated Data Can Poison Future AI Models

scientificamerican.com

71–80 of 87 posts

Re: AI-Generated Data Can Poison Future AI Models

#71

Earlier quoted context omitted.

That's far from unique to instagram. I loathe Stable Diffiusion and co solely because they've utterly FLOODED every cool art-adjacent website with endless mediocre derivative shit. Like there was always low-effort content of course, but holy fuck, there is SO MUCH MORE now. And some of these people are trying to CHARGE for this uninspired junk!!!

I agree with this despite using SD a lot myself. It's fun to use until you realize the majority of people posting stuff generated with it have almost no creativity, all generating the same things over and over again, mostly without any manual work involved. that uncanny realism style with the generic Stable Diffusion face and one of 5 different poses. The number of people putting any sort of effort into it is way, wa…

Honestly, if you want to make art, AI is a hindrance not a tool. You have to engineer what you say to maximum exactness, and even then it'll still just ignore certain words, or skip them over, or get basic details wrong.

Way back when all this stuff first popped off, I did try it out. I was unimpressed. It was like playing a game of telephone with my ideas, having to describe them into one end and have a thousand people repeat it to one another and make little contributions till it came out the other end, most of the time looking absolutely nothing like I expected.

People who say this makes art accessible... I dunno, I've never gotten it. I've seen people with all manner of disabilities, deformities, etc. all manage to express themselves creatively with practice and accessibility tools far more reliably than trying to make an AI pop out what you actually want. It seems to be the accessibility claims really only hold water if the accessibility feature is "I don't want to learn any skills" which... I mean, okay. But as with all art, your end product will reflect that level of care.

Re: AI-Generated Data Can Poison Future AI Models

#72
post #67

Earlier quoted context omitted.

That's far from unique to instagram. I loathe Stable Diffiusion and co solely because they've utterly FLOODED every cool art-adjacent website with endless mediocre derivative shit. Like there was always low-effort content of course, but holy fuck, there is SO MUCH MORE now. And some of these people are trying to CHARGE for this uninspired junk!!!

I definitely think the flooding of art spaces is hugely problematic, but it is pretty funny to watch people try to "be an artist" by putting essentially no effort in. It definitely points to a lack of understanding in the field when all these people are basically generating a ton of images that are all derived from the same models. There's a lack of understanding of supply and demand, when the expectation is that you…

I mean it's hard to fault those people when that was essentially how these were sold way back. "Automating art" and all that. All the most insufferable people on twitter jumping from shilling crypto scams to shilling AI and telling real artists in their ivory towers that their days were numbered.

Guess put it on the pile with all the other broken promises.

Re: AI-Generated Data Can Poison Future AI Models

#73

Earlier quoted context omitted.

This is exactly right. Model collapse does not exist in practice. In fact, LLMs trained on newer web scrapes have increased capabilities thanks to the generated output in their training data. For example, "base" pretrained models trained on scrapes which include generated outputs can 0-shot instruction follow and score higher on reasoning benchmarks. Intentionally produced synthetic training data takes this a step fu…

What happens if you train a model on nothing but AI-generated output, recursively? Does it eventually get inbred?

I want to observe that one of my favorite youtubers did exactly this with making the "uppest case" and "lowest case" letters.

https://www.youtube.com/watch?v=HLRdruqQfRk

I love this guy so much and wish he made far more videos.

Re: AI-Generated Data Can Poison Future AI Models

#74

Earlier quoted context omitted.

This is exactly right. Model collapse does not exist in practice. In fact, LLMs trained on newer web scrapes have increased capabilities thanks to the generated output in their training data. For example, "base" pretrained models trained on scrapes which include generated outputs can 0-shot instruction follow and score higher on reasoning benchmarks. Intentionally produced synthetic training data takes this a step fu…

What happens if you train a model on nothing but AI-generated output, recursively? Does it eventually get inbred?

Depends how good the AI output is, just like it depends how good the natural output is.

If most of it is bad but you can get a better AI to tag it as bad, then it's not necessarily a problem.

Re: AI-Generated Data Can Poison Future AI Models

#76
post #43

Computers need to be able to learn from the world at large, not just their own output. World models are needed to make progress.

There's no such thing as a world model. People do not have world models.

This is a confused term made up by 70s AI researchers, who had the continual problem that they didn't know any philosophy and kept making up their own metaphors for how intelligence might work, and then deciding that because they'd made it up it must be true, and also that if they wrote a computer program that had the same metaphors it must work.

"World model" just vaguely points at something people might do and assumes that if you make up a new thing it vaguely points at it'd help.

Re: AI-Generated Data Can Poison Future AI Models

#77

I'm not sure how much of a risk this is to LLMs in particular, but I feel like we're already seeing the impact on image AI models. Even though they're getting better at generating hands that make sense and other fine details, you can generally tell that an image is AI generated because it has a certain "style". Can't help but wonder if this is partly due to generated images contaminating the training data and causing…

It's because the models don't have an optimal aesthetic policy. Which would be difficult, but if they did have one, it wouldn't matter how much bad input data you added during pretraining.

Re: AI-Generated Data Can Poison Future AI Models

#78
post #34

Earlier quoted context omitted.

> I don't think a human mind would be improving if they're in a echo-chamber with no new information If this were true of humans, we would have never made it this far Humans are very capable of looking around themselves and thinking "I can do better than this", and then trying to come up with ways how LLMs are not

> Humans are very capable of looking around themselves and thinking "I can do better than this" Doesn't this require at least some perspective of what "better than this" means, which you could only know with at least a bit of outside influence in one way or another?

Every human has feelings and instincts, they answer what "better than this" means.

Yes, even in math and science, those were built on top of our feeling of "better than this" iterated over thousands of years.

Re: AI-Generated Data Can Poison Future AI Models

#79
post #43

Computers need to be able to learn from the world at large, not just their own output. World models are needed to make progress.

There's no such thing as a world model. People do not have world models. This is a confused term made up by 70s AI researchers, who had the continual problem that they didn't know any philosophy and kept making up their own metaphors for how intelligence might work, and then deciding that because they'd made it up it must be true, and also that if they wrote a computer program that had the same metaphors it must work…

On what basis are you saying that? I have a model in my head mapping out the world around me, so I know where things are etc and what I can do with all those things. How is that not a world model? Are you using a very strange definition of "world model"?

Re: AI-Generated Data Can Poison Future AI Models

#80
post #79

Earlier quoted context omitted.

There's no such thing as a world model. People do not have world models. This is a confused term made up by 70s AI researchers, who had the continual problem that they didn't know any philosophy and kept making up their own metaphors for how intelligence might work, and then deciding that because they'd made it up it must be true, and also that if they wrote a computer program that had the same metaphors it must work…

On what basis are you saying that? I have a model in my head mapping out the world around me, so I know where things are etc and what I can do with all those things. How is that not a world model? Are you using a very strange definition of "world model"?

> On what basis are you saying that?

This is a longstanding critique of GOFAI; see Hubert Dreyfuss and Phil Agre.

https://pages.gseis.ucla.edu/faculty/agre/critical.html

> I have a model in my head mapping out the world around me, so I know where things are etc and what I can do with all those things.

No you don't; a map is not the territory, and is necessarily wrong, which means that if you had such a model and were actually relying on it you wouldn't be able to do things you obviously can do in real life.

You have an inaccurate memory of the world and you update it, only as much as you need to[0], as you go, in order to do a specific task.

[0] probably a little less than you need to, because you want to save thinking energy

Post reply on HN