Training open-source LLMs on ChatGPT output is a really bad idea.
1–10 of 78 posts
Re: Training open-source LLMs on ChatGPT output is a really bad idea.
#2Re: Training open-source LLMs on ChatGPT output is a really bad idea.
#3A subset will give you a subset of the knowledge, it’s no free lunch
Re: Training open-source LLMs on ChatGPT output is a really bad idea.
#4Re: Training open-source LLMs on ChatGPT output is a really bad idea.
#5It won’t work because will need as much training data as ChatGPT to get to its general knowledge level. A subset will give you a subset of the knowledge, it’s no free lunch
Some recent research hash showed reasonable success transferring fine tuning between models using outputs.
Re: Training open-source LLMs on ChatGPT output is a really bad idea.
#6Re: Training open-source LLMs on ChatGPT output is a really bad idea.
#7I wonder how OpenAI are going to avoid the problem after the web is littered with its content?
Re: Training open-source LLMs on ChatGPT output is a really bad idea.
#8To be fair, we're already there, and we've been there for at least 10 years now. I'd wager >75% of the internet is garbage: auto-generated blog posts, programmatically-permuted ads, YouTube videos that mainly regurgitate other sources. Email is mostly garbage and the only reason it's usable is because spam filters have gotten pretty good. Even non-trivial amounts of heavily-curated social media (Twitter/FB/IG) is purely spam.
Re: Training open-source LLMs on ChatGPT output is a really bad idea.
#9I wonder how OpenAI are going to avoid the problem after the web is littered with its content?
Re: Training open-source LLMs on ChatGPT output is a really bad idea.
#10I wonder how OpenAI are going to avoid the problem after the web is littered with its content?