Live data from Hacker News

The False Promise of Imitating Proprietary LLMs

arxiv.org

31–40 of 90 posts

Re: The False Promise of Imitating Proprietary LLMs

#31
post #13

From the Conclusion: "Finally, our work raises ethical and legal questions, including whether the open-source community should continue to advance progress by “stealing” what OpenAI and other companies have done, as well as what legal countermeasures companies can take to protect and license intellectual property." Really???

their work didn't they did.

I'm going to need verifiable proof this wasn't written by chatGPT as propaganda.

Re: The False Promise of Imitating Proprietary LLMs

#33
"Second, given the large gap between LLaMA and ChatGPT (the latter model is faster, cheaper, and more accurate), "

No it's not, llama would be cheaper and likely faster if you ran it on the same scale, actually there've been a few calcs done, that running llama 65b if you're at 100% usage is cheaper than 3.5turbo per token. Also comparing them for accuracy isn't fair comparison, one is a foundational model, one is an instruct tuned model. Perhaps compare llama 65b with gpt3.

Re: The False Promise of Imitating Proprietary LLMs

#34
post #2

The authors conduct automated, more methodical evaluations of LLMs finetuned to imitate ChatGPT outputs, and find that, despite superficial/informal appearances to the contrary, the base LLMs close little to none of the gap to ChatGPT on tasks that are not heavily supported in the imitation data. It's not good news for the open LLM ecosystem.

This is a very weird type of paper. They take a specific approach, then make arguments about a broad class of approaches that are under constant development. The finding that distilled LLMs must be more specialized than the giant LLMs that train them is unsurprising; nobody at this point expects a 13B parameter model to succeed with the same accuracy at the broad range of tasks supported by what may be a 1T parameter…

Even if they don't start out expecting that, people might be fooled by how it behaves when they try it out. So it seems useful to point out that initial impressions based on crowdsourced evaluations are misleading:

> We were initially surprised by how much imitation models improve over their base models: they are far better at following instructions, and their outputs appear similar to ChatGPT’s. This was further supported by both human and GPT-4 evaluations, where the outputs of our best imitation model were rated as competitive with ChatGPT (e.g., Figure 1, left).

("Competitive" meaning that 70% outputs seemed about as good.)

Re: The False Promise of Imitating Proprietary LLMs

#36
Sensational title that misrepresents the message in paper.

However, when conducting more targeted automatic evaluations, we found that the imitation models close little to none of the large gap between LLaMA and ChatGPT. In particular, we demonstrate that imitation models improve on evaluation tasks that are heavily supported in the imitation training data. On the other hand, the models do not improve (or even decline in accuracy) on evaluation datasets for which there is little support. For example, training on 100k ChatGPT outputs from broad-coverage user inputs provides no benefits to Natural Questions accuracy (e.g., Figure 1, center), but training exclusively on ChatGPT responses for Natural-Questions-like queries drastically improves task accuracy.

Just because this might not be the way to replicate the performance of ChatGPT across all tasks, it seems to work quite well on whichever tasks are in the imitation learning. That is still a big win.

Later on this also works for factual correctness. (leaving aside the argument whether this is the right approach for factuality)

For example, training on 100k ChatGPT outputs from broad-coverage user inputs provides no benefits to Natural Questions accuracy (e.g., Figure 1, center), but training exclusively on ChatGPT responses for Natural-Questions-like queries drastically improves task accuracy.

Re: The False Promise of Imitating Proprietary LLMs

#37
post #28

Earlier quoted context omitted.

Like a torrent of the last GoT season then? … with compression.

Imagine the GoT producers used GRRM's books without licensing and then claim copyright on the series. Does OpenAI have the rights on all the texts they used to train their GPTs?

i think we agree

Re: The False Promise of Imitating Proprietary LLMs

#38
post #36

Sensational title that misrepresents the message in paper. However, when conducting more targeted automatic evaluations, we found that the imitation models close little to none of the large gap between LLaMA and ChatGPT. In particular, we demonstrate that imitation models improve on evaluation tasks that are heavily supported in the imitation training data. On the other hand, the models do not improve (or even declin…

To be fair, this paper has been made obsolete in its entirety with recent research. It's not really their fault, but folks need to start publishing faster as posters or something if they want to provide something relevant.

A better title, knowing what we now, might be "To outperform GPT4, do more than imitating"

Re: The False Promise of Imitating Proprietary LLMs

#39
post #36

Sensational title that misrepresents the message in paper. However, when conducting more targeted automatic evaluations, we found that the imitation models close little to none of the large gap between LLaMA and ChatGPT. In particular, we demonstrate that imitation models improve on evaluation tasks that are heavily supported in the imitation training data. On the other hand, the models do not improve (or even declin…

To be fair, this paper has been made obsolete in its entirety with recent research. It's not really their fault, but folks need to start publishing faster as posters or something if they want to provide something relevant. A better title, knowing what we now, might be "To outperform GPT4, do more than imitating"

Link to said research?

Re: The False Promise of Imitating Proprietary LLMs

#40

This is exactly the reason OpenAI isn't afraid of the open-source community, like many kneejerk opponents of regulatory capture assume (they are probably still afraid of Google). Also why they still do the expensive and cumbersome RLHF training, instead of those deceptively cheap and fast finetunes. They understand their own tech and why there isn't free lunch. Recently, John Schulman explained the issue with behavio…

Before the internet I use to laugh my ass of at school friends who would wonder/debate about something but never bother to look it up, in stead they had elaborate collective "hallucinations", they imagined the facts until they were satisfied their answer was well reasoned enough, then they would consider it a fact. We all do this at times (at all ages) but one must learn to shut down the train of thought, stop polluting your memory.

I remember one from very early in life. I postulated out loud that Jerusalem, being the birth place of Jesus, must be the most peaceful place on earth. All those loving and caring religious people who work so hard to be good people. That their religions are slightly different shouldn't matter to Jesus message?

That the LLM's can consume such huge amounts of data doesn't mean they matured beyond that rather infantile mind set.

In the video you linked he explained that training it to learn to say it doesn't know will trigger false negatives.

The correct formula I imagine (hah!) is to wonder if the question is of interest to the model and to ask someone else for answers or some help figuring out the question. The human will just have to wait.

What is completely hilarious to me is that we all have heads full of learned answers for which we have no idea "why it is so" or at the very least lack that what would have one arrive at that solution. I get what Archimedes realized in the bathtub but what I want is the mechanism by which he arrived at such wild ideas. Could it be that learning a lot of facts would be the exact opposite kind of conditioning?

My mind now says this must be why we humans expire so fast. You keep the calcified brains around for a while as data storage but the focus of project humanity must be to create young ones. I will have to ponder this fact free line of reasoning some more. Perhaps I will find ways to convince myself I know something.

It is a fun thought that people created AI, we really want to believe we did. If enough pretend it is true no one can take it away from us.

If you want people to think you are intelligent you tell them things they already know and hide your sources.

Post reply on HN