Live data from Hacker News

The False Promise of Imitating Proprietary LLMs

arxiv.org

21–30 of 90 posts

Re: The False Promise of Imitating Proprietary LLMs

#22
post #2

The authors conduct automated, more methodical evaluations of LLMs finetuned to imitate ChatGPT outputs, and find that, despite superficial/informal appearances to the contrary, the base LLMs close little to none of the gap to ChatGPT on tasks that are not heavily supported in the imitation data. It's not good news for the open LLM ecosystem.

This is a very weird type of paper. They take a specific approach, then make arguments about a broad class of approaches that are under constant development. The finding that distilled LLMs must be more specialized than the giant LLMs that train them is unsurprising; nobody at this point expects a 13B parameter model to succeed with the same accuracy at the broad range of tasks supported by what may be a 1T parameter…

> nobody at this point expects a 13B parameter model to succeed with the same accuracy at the broad range of tasks supported by what may be a 1T parameter model

I think a lot of people believe exactly that. To take one example from the "We Have No Moat" essay:

"It doesn’t take long before the cumulative effect of all of these fine-tunings overcomes starting off at a size disadvantage. Indeed, in terms of engineer-hours, the pace of improvement from these models vastly outstrips what we can do with our largest variants, and the best are already largely indistinguishable from ChatGPT." - https://www.semianalysis.com/p/google-we-have-no-moat-and-ne...

Re: The False Promise of Imitating Proprietary LLMs

#23

The breathtaking audacity of calling distilling GPT4 'stealing' when GPT4 trained on data it has no proprietary right to.

Just because someone can convert text to numbers doesn’t mean they have a right to the numbers. That’s like trying to own the emotion a book has on someone, or the things they see in mind when they read it.

What I find rather amusing is they spend the whole paper dismissing it as ineffective yet still feel the need to worry about the 'ethics' and 'legality'. They don't cite anything with regards to a discussion/evidence of either, of course, and looking at the authorship list I don't believe any of them are lawyers or ethics experts.

Re: The False Promise of Imitating Proprietary LLMs

#24

The breathtaking audacity of calling distilling GPT4 'stealing' when GPT4 trained on data it has no proprietary right to.

Just because someone can convert text to numbers doesn’t mean they have a right to the numbers. That’s like trying to own the emotion a book has on someone, or the things they see in mind when they read it.

[flagged]

Re: The False Promise of Imitating Proprietary LLMs

#25

The breathtaking audacity of calling distilling GPT4 'stealing' when GPT4 trained on data it has no proprietary right to.

Just because someone can convert text to numbers doesn’t mean they have a right to the numbers. That’s like trying to own the emotion a book has on someone, or the things they see in mind when they read it.

Like a torrent of the last GoT season then?

… with compression.

Re: The False Promise of Imitating Proprietary LLMs

#26
post #13

From the Conclusion: "Finally, our work raises ethical and legal questions, including whether the open-source community should continue to advance progress by “stealing” what OpenAI and other companies have done, as well as what legal countermeasures companies can take to protect and license intellectual property." Really???

I think the creators of all the scraped training data would like to talk about intellectual property too

Re: The False Promise of Imitating Proprietary LLMs

#27
post #19

If they really didn't test anything bigger than 13b, as their abstract states, then this doesn't even seem worth reading through.

The "Google has no moat" thing claimed that Vicuna-13B was almost as good as ChatGPT and this paper seemingly refutes that.

Claims made in a leaked blog post shouldn't be considered as having any sort of scientific authority. That whole "no moat" piece has exactly the tone I would expect from an over-confident Googler who has essentially been following all of this by watching various Discord channels and browsing hacker news. That isn't how science is done. It shouldn't be how business is done, but people seem to really enjoy these everything-is-actually-simple narratives.

Re: The False Promise of Imitating Proprietary LLMs

#28

Earlier quoted context omitted.

Just because someone can convert text to numbers doesn’t mean they have a right to the numbers. That’s like trying to own the emotion a book has on someone, or the things they see in mind when they read it.

Like a torrent of the last GoT season then? … with compression.

Imagine the GoT producers used GRRM's books without licensing and then claim copyright on the series.

Does OpenAI have the rights on all the texts they used to train their GPTs?

Re: The False Promise of Imitating Proprietary LLMs

#29
post #2

The authors conduct automated, more methodical evaluations of LLMs finetuned to imitate ChatGPT outputs, and find that, despite superficial/informal appearances to the contrary, the base LLMs close little to none of the gap to ChatGPT on tasks that are not heavily supported in the imitation data. It's not good news for the open LLM ecosystem.

This is a very weird type of paper. They take a specific approach, then make arguments about a broad class of approaches that are under constant development. The finding that distilled LLMs must be more specialized than the giant LLMs that train them is unsurprising; nobody at this point expects a 13B parameter model to succeed with the same accuracy at the broad range of tasks supported by what may be a 1T parameter…

[deleted]

Re: The False Promise of Imitating Proprietary LLMs

#30

The breathtaking audacity of calling distilling GPT4 'stealing' when GPT4 trained on data it has no proprietary right to.

Just because someone can convert text to numbers doesn’t mean they have a right to the numbers. That’s like trying to own the emotion a book has on someone, or the things they see in mind when they read it.

I would like the big players to argue that they have some right to the numbers as it has important applications to BitTorrent and cryptography too for that matter.
Post reply on HN