Live data from Hacker News

AI-Generated Data Can Poison Future AI Models

scientificamerican.com

11–20 of 87 posts

Re: AI-Generated Data Can Poison Future AI Models

#12
post #8

This reminds me of how fascinated I was as a kid of the artifacts you get from recursively photocopying a piece of paper.

I watched someone in the printer room at the computer science department gradually photocopy from white to black, and back again, over the span of 300 pieces of paper, by altering the thresholds of the photocopyer.

They didn’t graduate to become computer scientists, but did indeed get admitted to the royal school of art the year after.

I found it strangely therapeutic.

Re: AI-Generated Data Can Poison Future AI Models

#13
post #4

I think it's interesting that human minds generally (though not always!) improve when exposed to the output of other human minds. It seems to be the opposite for current LLMs.

A more appropriate analogy would be isolating someone from the rest of the world and only being able to read their own writings from now on.

While some persons can strive in these kind of environment (think Kant for example), many would become crazy.

Re: AI-Generated Data Can Poison Future AI Models

#14
post #4

I think it's interesting that human minds generally (though not always!) improve when exposed to the output of other human minds. It seems to be the opposite for current LLMs.

humans haven’t been had the same set of all encompassing “training experiences” like LLMs have. we each a subset of knowledge that may overlap with some other’s knowledge, but is largely unique. so when we interact with each other we can learn new things, but with LLMs I imagine it is a group of experienced but antiquated professors developing their own set of out of touch ideas

Re: AI-Generated Data Can Poison Future AI Models

#17
post #5

Unless the internet is no longer useful because there is no way to find anything reliable, there would be enough signal to train and align models.

I question whether it'll matter. There is so much language data already, unlocking a little more isn't going to be the difference maker for AGI.

Re: AI-Generated Data Can Poison Future AI Models

#18
post #6
post #4

I think it's interesting that human minds generally (though not always!) improve when exposed to the output of other human minds. It seems to be the opposite for current LLMs.

Maybe it's less about "Human VS Robot" and more about exposure to "Original thoughts VS mass-produced average thoughts". I don't think a human mind would be improving if they're in a echo-chamber with no new information. I think the reason the human mind is improving is because we're exposed to new, original and/or different thoughts, that we hadn't considered or come across before. Meanwhile, a LLM will just regurgi…

> I don't think a human mind would be improving if they're in a echo-chamber with no new information

If this were true of humans, we would have never made it this far

Humans are very capable of looking around themselves and thinking "I can do better than this", and then trying to come up with ways how

LLMs are not

Re: AI-Generated Data Can Poison Future AI Models

#20

Some perspectives from someone working in the image space. These tests don't feel practical - That is, they seem intended to collapse the model, not demonstrate "in the wild" performance. The assumption is that all content is black or white - AI or not AI - and that you treat all content as equally worth retraining on. It offers no room for assumptions around data augmentation, human-guided quality discrimination, or…

> Use the model to generate some AI output. Then use that output to train a new instance of the model and use the resulting output to train a third version, and so forth. With each iteration, errors build atop one another. The 10th model, prompted to write about historical English architecture, spews out gibberish about jackrabbits.

That this happens doesn't surprise me, but I'd love to see a curve of how each organic vs machine content mixe ratio results in model collapse over N generations.

Post reply on HN