The breathtaking audacity of calling distilling GPT4 'stealing' when GPT4 trained on data it has no proprietary right to.
The False Promise of Imitating Proprietary LLMs
71–80 of 90 posts
Re: The False Promise of Imitating Proprietary LLMs
#72Conspiracy theory: Is that the reason why GPT-4 is not available as an API? So people wouldn't siphon off it's capabilities?
1. It is available via API 2. Likely not a conspiracy theory. Newer models don't have logits available and that's almost certainly because they didn't want other labs distilling from them.
This definitely is a smoking gun. Keeping the crown jewels for themselves.
Re: The False Promise of Imitating Proprietary LLMs
#73This is exactly the reason OpenAI isn't afraid of the open-source community, like many kneejerk opponents of regulatory capture assume (they are probably still afraid of Google). Also why they still do the expensive and cumbersome RLHF training, instead of those deceptively cheap and fast finetunes. They understand their own tech and why there isn't free lunch. Recently, John Schulman explained the issue with behavio…
Before the internet I use to laugh my ass of at school friends who would wonder/debate about something but never bother to look it up, in stead they had elaborate collective "hallucinations", they imagined the facts until they were satisfied their answer was well reasoned enough, then they would consider it a fact. We all do this at times (at all ages) but one must learn to shut down the train of thought, stop pollut…
Most humans are going to outlast whatever they produced during their lives. If anything, human bodies are among the most durable "goods" in the economy. Only real estate, public infrastructure and recorded knowledge (including genes) last longer than a human lifetime. How many of the things you buy and own are going to outlast you?
Re: The False Promise of Imitating Proprietary LLMs
#74Particularly this statement seems relevant: "We provide a detailed analysis of chatbot performance based on both human and GPT-4 evaluations showing that GPT-4 evaluations are a cheap and reasonable alternative to human evaluation. Furthermore, we find that current chatbot benchmarks are not trustworthy to accurately evaluate the performance levels of chatbots. A lemon-picked analysis demonstrates where Guanaco fails compared to ChatGPT."
Re: The False Promise of Imitating Proprietary LLMs
#75Earlier quoted context omitted.
Before the internet I use to laugh my ass of at school friends who would wonder/debate about something but never bother to look it up, in stead they had elaborate collective "hallucinations", they imagined the facts until they were satisfied their answer was well reasoned enough, then they would consider it a fact. We all do this at times (at all ages) but one must learn to shut down the train of thought, stop pollut…
Humans don't expire very quickly. Most humans are going to outlast whatever they produced during their lives. If anything, human bodies are among the most durable "goods" in the economy. Only real estate, public infrastructure and recorded knowledge (including genes) last longer than a human lifetime. How many of the things you buy and own are going to outlast you?
All the goods we produce are designed to last for a specific time. We can easily make them more durable and with some serious effort they could last longer than we can imagine. It would be expensive, it might be beneficial but who wants to pay for benefits 100 or 200 years into the future?
Re: The False Promise of Imitating Proprietary LLMs
#76The breathtaking audacity of calling distilling GPT4 'stealing' when GPT4 trained on data it has no proprietary right to.
Re: The False Promise of Imitating Proprietary LLMs
#77Earlier quoted context omitted.
And neither is anyone who has played with these new LLMs, found them so-so, and wondered whether the hype was warranted.
I’m curios as to why you think the hype isn’t warranted. If you go through my history (you don’t have too I’ll sum it up), you’ll see that I’m not impressed by the capabilities of LLM to actually do my work. Not for a lack of trying, but because ChatGPT simply tells too many lies. We’ve yet to get it to really do anything that wasn’t fairly basic, or solved a billion times on the internet anyway. Similarly we’ve stop…
That said, the rest of your comment is spot-on.
Paul G says this too that ChatGPT expertise is the same as a journalist's expertise. Its output seems impressive until it is on a subject you know very well.
GPT-x is like a wide-eyed intern or junior team member who loves to shoot its mouth because it has been told to be assertive and vocal. The good thing is that it is willing to learn.
Now, if this is true of GPT-x which is pretty much the benchmark against which every open source LLM is being measured, you can guess for yourself how much room these open source LLMs still have to cover.
Re: The False Promise of Imitating Proprietary LLMs
#78The breathtaking audacity of calling distilling GPT4 'stealing' when GPT4 trained on data it has no proprietary right to.
Was GPT-4 trained on data that was acquired illegally? Or was it trained on data acquired legally that OpenAI didn't have the rights to redistribute? There is a difference. In the latter case, whether it counts as "stealing" would come down to whether or not GPT-4 counts as a derivative work, or some similar legal concept.
Scribd has lots of pdfs of books that are copyrighted. The Washington Post article mentions there are several other places it downloaded and scraped pdfs of copyrighted textbooks, etc
Re: The False Promise of Imitating Proprietary LLMs
#79From the Conclusion: "Finally, our work raises ethical and legal questions, including whether the open-source community should continue to advance progress by “stealing” what OpenAI and other companies have done, as well as what legal countermeasures companies can take to protect and license intellectual property." Really???
Re: The False Promise of Imitating Proprietary LLMs
#80The jump between llama 13B and 30B is quite significant. And their instruction finetuning is not SOTA I don't think, though the point about general knowledge is a good one: instruction llama lies very confidently. But one great thing about open source LLMs is that you can specialize them in various tasks with affordable LORA training, enough to easily beat GPT4 in a specific niche.