Live data from Hacker News

AI models collapse when trained on recursively generated data

nature.com

111–120 of 212 posts

Re: AI models collapse when trained on recursively generated data

#111
post #90

Earlier quoted context omitted.

I think this is it. Generated data is ok if you're curating it to make sure nothing bad, wrong or insensible comes in. Basically still needs a human in the loop.

No, bad/wrong/nonsense is not the only risk here. You're missing the main point that the authors are making: the shape of the distribution gets changed by this process. A model trained on human data will produce fewer high-perplexity examples than it was trained on (you can see this in Fig 1b, even between generation 0 and 1). In a literal information theory sense, these perplexity values indicate how much informatio…

LLMs are milking us of knowledge and skills, repackage them and give it back to us. Models interact with the internet, humans and code execution. They are exploring. Lots of exploring now happens in the chat room, a place where ideas are first tried out. With billions of users, the volume of information LLMs collect from us is huge. We bring references, guidance and feedback right into its mouth, the LLM doesn't even need to do anything like crawling.

Imagine how many things we know, things we accumulated in our life experience, that were never written down anywhere. That information was lost to others. But now we use LLM assistants, so they get to be in the loop and collect tidbits of human life experience that is not written on the internet. And soon they will also work on audio/video and travel with us everywhere, seeing what we show them.

Re: AI models collapse when trained on recursively generated data

#112
post #95

I must be missing something. Training on the output of your system as if it were validated input seems like an obvious no-no. I'm not talking about using synthetic data (however that might be created in this situation), but rather using anything and everything found on the web as if it were "real", i.e. as if it were human-generated texts rather than the output of the LLM. In this case of course there are multiple LL…

It's _relatively_ easy, I think to filter out sites with a large proportion of low quality ai-generated glurge.

Then you're left with a lot of AI generated or assisted content that has quite often been filtered and modified by humans, so that might mitigate some of the problems that cause model collapse because the filtered content _should_ better reflect reality or desirable output?

Re: AI models collapse when trained on recursively generated data

#113
post #95

I must be missing something. Training on the output of your system as if it were validated input seems like an obvious no-no. I'm not talking about using synthetic data (however that might be created in this situation), but rather using anything and everything found on the web as if it were "real", i.e. as if it were human-generated texts rather than the output of the LLM. In this case of course there are multiple LL…

You would think so, but people like Sam Altman have suggested that they can use AI-generated data to train their own models. See here: https://www.nytimes.com/2024/04/06/technology/tech-giants-ha...

Training on ai-generated data isn't a problem, and has been routinely done by everyone for 18 mo +.

The issue is training on 'indiscriminate' ai-generated data. This just leads to more and more degenerate results. No one is doing this however, there is always some kind of filtering to select which generated data to use for training. So the finding of that paper are entirely not surprising, and frankly, intuitive and already well known.

Re: AI models collapse when trained on recursively generated data

#114
post #97
post #23

> We find that indiscriminate use of model-generated content in training causes irreversible defects in the resulting models The key word there is "indiscriminate". All of the big AI labs have been training on synthetic data for at least a year at this point, but they're doing so deliberately. I don't think the "model collapse" problem is particularly important these days. The people training models seem to have that…

The question (which I raised in a top-level comment before reading your post) is whether there is any such thing as "discriminate" use of web data. Synthetic data created in the same lab as the LLM is discriminate, but what the authors of the paper are saying (if I read it correctly) is that scraping the web is not currently done in a discriminate way. And it's not at all clear to me that there is a discriminate way…

I get the impression that scraping the web isn't nearly as important a source of LLM training data as it used to be.

Everyone is trimming down their training data based on quality - there are plenty of hints about that in the Llama 3.1 paper and Mistral Large 2 announcement.

OpenAI are licensing data from sources like the Associated Press.

Andrej Karpathy said this: https://twitter.com/karpathy/status/1797313173449764933

> Turns out that LLMs learn a lot better and faster from educational content as well. This is partly because the average Common Crawl article (internet pages) is not of very high value and distracts the training, packing in too much irrelevant information. The average webpage on the internet is so random and terrible it's not even clear how prior LLMs learn anything at all.

Re: AI models collapse when trained on recursively generated data

#115
post #99

Earlier quoted context omitted.

Yeah, I raised the same issue before reading your post; ninja'd I am. I like your "cheese and chalk".

I always preferred sugar and shit. Obviously, that is profane. But, I consider profanity be seen as the part of speech it really is.

Profane is, if you will, fucking fine.

But "cheese and chalk" is a great analogy because both are sources of calcium, but cheese is much better for the human body. It carries useful info.

Re: AI models collapse when trained on recursively generated data

#116
post #61

The article contains no proof of theorem 3.1 and finding counterexamples seems trivial. Adult male weight can be modeled by N(85, 20). You can recursively "train" the model on data it generates without having it collapse. It will stay stationary as long as the samples are large enough.

I believe that counterexample only works in the limit where the sample size goes to infinity. Every finite sample will have μ≠0 almost surely.(Of course μ will still tend to be very close to 0 for large samples, but still slightly off) So this means the sequence of μₙ will perform a kind of random walk that can stray arbitrarily far from 0 and is almost sure to eventually do so.

Fair point about the mean, but I don't see how the random walk causes the standard deviation to shrink towards zero.

Re: AI models collapse when trained on recursively generated data

#117
post #14

Is this an artifact of floating point precision or a fundamental mathematical truth.

It’s a lossy transformation, so you’re losing information each time. It’s never going to add information. However, some information is junk that obscures the good stuff. It’s likely that how they train today is very inefficient compared to what’s possible, and there will be smarter ways to transform preexisting data so that it’s a better dataset to train on, without losing very much. Papers like this one show what no…

> there will be smarter ways to transform preexisting data so that it’s a better dataset to train on, without losing very much

Like, take for example search. Instead of training on a bunch of scraped texts, you take one prompt, select 10 references, and use it to synthesize an answer. Referencing multiple texts gives you more than training on them directly. The LLM could catch contradictions, observe the distribution of human opinions, note if the topic is controversial. And then output a wikipedia-like article. Do this billions of times, and you got a refined dataset. You can iterate on top, using the articles as source and writing meta articles. Or just silly studies like writing a paper about "Characters named Charlie in literature". You can slice and dice the data in any way, and analyze the cross section.

Re: AI models collapse when trained on recursively generated data

#118
post #6
post #2

Which is good background to this story about Reddit locking down robots.txt and trying to get money from the AI teams scraping their content. https://news.ycombinator.com/item?id=41057033

If they're considering Reddit content to be free of generated material, I've got bad news for them. It's not quite the Chernobyl-grade hole that Pinterest has become, but it's hardly "low background".

I still believe reddit is an amazing source. Any article you read on reddit, chances are the comments are better than the original text. They will debunk the article, present a diversity of reactions, and most importantly, they will be grounded in public opinion unlike the press which caters to money interests.

You just copy-paste a conversation into the LLM and ask for an article. For taste, here is one generated from this very conversation. https://pastebin.com/raw/JFH6PGqg

Re: AI models collapse when trained on recursively generated data

#119
post #22

Earlier quoted context omitted.

I'm long on synthetic data. If you think about evolution and hill climbing, of course it works. You have a pool of information and you accumulate new rearrangements of that information. Fitness selects for the best features within the new pool of data (For primates, opposable thumbs. For AI art, hands that aren't deformed.) It will naturally drift to better optima. RLHF, synthetic data, and enrichment are all we need…

It's important to climb the right hill.

And to have a very well-tuned sense up vs down when the hill is almost flat...

Re: AI models collapse when trained on recursively generated data

#120
post #23

> We find that indiscriminate use of model-generated content in training causes irreversible defects in the resulting models The key word there is "indiscriminate". All of the big AI labs have been training on synthetic data for at least a year at this point, but they're doing so deliberately. I don't think the "model collapse" problem is particularly important these days. The people training models seem to have that…

They make it clear in the paper that their primary "real-world" concern is that it's difficult to distinguish synthetic data from real human interaction when scraping data from the web. This will only get worse over time with our current way of doing things.

How are they supposed to deliberately train on synthetic data when they don't know whether it is (synthetic) or not?

Also, do you not feel that it is presumptuous to dismiss a body of work in a few sentences with a "seems fine to me"?

Post reply on HN