Live data from Hacker News

The False Promise of Imitating Proprietary LLMs

arxiv.org

61–70 of 90 posts

Re: The False Promise of Imitating Proprietary LLMs

#61

Earlier quoted context omitted.

By that logic, you also need to accept that no one should ever need to pay you for creating artifacts that are not bound to the physical world solely. I assume you work for free for your employer or in a space that is not “dealing” with data, information, bits, whatsoever.

Hairdressers charge for a service and none of them will assume that they "own" your hair.

They are transforming physical objects very much the same way a carpenter does. The service industry is not equal to Tech / digital. A hairdresser does not create Bits or data. I would also argue in this particular case you are wrong. You hand over the hair on the ground to them which they then dispose or maybe resell (maybe without explicit consent but at least implicit). If that wasn’t the case they would commit theft when they dispose your hair…

Re: The False Promise of Imitating Proprietary LLMs

#62
post #60

Earlier quoted context omitted.

Hairdressers charge for a service and none of them will assume that they "own" your hair.

I am honestly shocked that this hasn't happened, what with how the world has been going in recent decades.

Well, they probably can since you give them consent to keep your hair when you leave the shop… Disposing it would otherwise be considered theft, no?

Re: The False Promise of Imitating Proprietary LLMs

#64

Conspiracy theory: Is that the reason why GPT-4 is not available as an API? So people wouldn't siphon off it's capabilities?

1. It is available via API 2. Likely not a conspiracy theory. Newer models don't have logits available and that's almost certainly because they didn't want other labs distilling from them.

[deleted]

Re: The False Promise of Imitating Proprietary LLMs

#66
post #26
post #13

From the Conclusion: "Finally, our work raises ethical and legal questions, including whether the open-source community should continue to advance progress by “stealing” what OpenAI and other companies have done, as well as what legal countermeasures companies can take to protect and license intellectual property." Really???

I think the creators of all the scraped training data would like to talk about intellectual property too

I've written hundreds of thousands of words, possibly millions on various sites including HN, Reddit, blogs, Stackoverflow, forgotten platforms, etc. Doesn't seem right that OpenAI/LLMs can use my intellect but I can't use theirs.

Re: The False Promise of Imitating Proprietary LLMs

#67
post #22

Earlier quoted context omitted.

This is a very weird type of paper. They take a specific approach, then make arguments about a broad class of approaches that are under constant development. The finding that distilled LLMs must be more specialized than the giant LLMs that train them is unsurprising; nobody at this point expects a 13B parameter model to succeed with the same accuracy at the broad range of tasks supported by what may be a 1T parameter…

> nobody at this point expects a 13B parameter model to succeed with the same accuracy at the broad range of tasks supported by what may be a 1T parameter model I think a lot of people believe exactly that. To take one example from the "We Have No Moat" essay: "It doesn’t take long before the cumulative effect of all of these fine-tunings overcomes starting off at a size disadvantage. Indeed, in terms of engineer-hou…

That essay works in a context of specific datasets and tasks, which are referenced in the surrounding sentences and paragraphs. They are saying that for a particular "emergent" capability you might reach with a giant LLM, you might get there more efficiently with distillation / LoRa.

My comment is about generality, which is the remaining advantage of giant models.

Re: The False Promise of Imitating Proprietary LLMs

#68
post #53

Earlier quoted context omitted.

This is a very weird type of paper. They take a specific approach, then make arguments about a broad class of approaches that are under constant development. The finding that distilled LLMs must be more specialized than the giant LLMs that train them is unsurprising; nobody at this point expects a 13B parameter model to succeed with the same accuracy at the broad range of tasks supported by what may be a 1T parameter…

That is exactly what people are expecting, and largely because of misleading metrics thrown around to claim ridiculous things like e.g. Vicuna-13b being nearly as good as GPT-3.5. It even shows up in the comments here, and if you go to any tangentially related subreddit, that's the kind of stuff that gets told as "everybody knows" to people setting up a local LLM for the first time.

The Vicuna headline was def an overreach, although the main text admits pretty readily that their performance test (asking ChatGPT to evaluate quality) is not rigorous [1]. I'm sure that has set a lot of people pontificating about AGI with tiny models, but I can't imagine anyone who has worked directly with fine tuning having that impression.

The comments I see here are not about that. They are about small models succeeding at specific tasks, which is affirmed by this paper. Most applications of LLMs are not general purpose chat bots, so this is not bad news for most of the distill/fine tune community.

[1] https://lmsys.org/blog/2023-03-30-vicuna/

Re: The False Promise of Imitating Proprietary LLMs

#69
post #13

From the Conclusion: "Finally, our work raises ethical and legal questions, including whether the open-source community should continue to advance progress by “stealing” what OpenAI and other companies have done, as well as what legal countermeasures companies can take to protect and license intellectual property." Really???

[deleted]

Re: The False Promise of Imitating Proprietary LLMs

#70
post #56

Earlier quoted context omitted.

Good news for alignment though. This gives me a tiny amount of hope.

So, LLMs aligned with the interests of our corporate overlords and that nebulous "national security" thing that somehow always translates to more surveillance and less due process?

This tech has about as much chance to continue unregulated as highly enriched uranium. There is no future-path that includes unregulated AI.

I don't like horrific government abuse of residents,and I would not mind throwing most billionaire CEOs into a pool of alligators and dissolving their corporations. I don't like Altman, I think he's a smart person with NOBUS-level reckless hubris who is softballing the magnitude of the dangera to wet. The status quo is not good and it's getting worse.z

It doesn't matter. 5 people with launch-all-the-nukes buttons is better than 500 million.

Post reply on HN