Live data from Hacker News

QLoRA: Efficient Finetuning of Quantized LLMs

arxiv.org

71–80 of 116 posts

Re: QLoRA: Efficient Finetuning of Quantized LLMs

#71
post #62
post #8

Earlier quoted context omitted.

Altman’s push for regulatory capture makes so much sense given how fast this field is going. Open models you can run on regular hardware are still behind GPT-4 by some distance but they are closing in at a rate that leads me to believe there’s not much of a moat there. There’s also a ton of promising work on quantization and pruning and other acceleration and compression techniques to make more powerful models run on…

>Open models... are still behind GPT-4 by some distance but they are closing in at a rate that leads me to believe there’s not much of a moat there Commercially, this doesn't matter even if true. As I keep reminding here: moats are rarely about raw tech. Moats are much more about integrations, brand ("Nobody got fired for buying X"), access to API/raw compute... If OpenAI have a superior Office 365 integration they'l…

I agree with you on that a good integrator can get successful, but we should not forget here that the underlying code running all of this is in the order of hundreds of lines (see alpaca.cpp for example) in combination with a boatload of money required for training (that is quickly going down), to simplify the argument a bit.

It's hardly comparable to the gigantic ecosystem of services and micro-services that your comparison alludes to with IBM and Microsoft. OpenAI is nowhere near such a brand recognition, and grassroots support for it within a company would quickly move to the next free chatgpt-clone that is better or cheaper or faster or more accessible just as what happened with Dalle-2.

The value of the models themselves will quickly come down to just a slight bit over the underlying hardware costs and Altman knows this.

Microsoft bought themselves a $10B time window to try to do what you're saying, so let's see :) But even for them, when they've built LLM-adaptations to their most popular products, it's fairly simple to just swap it out with something new and more shiny and cheaper that's not OpenAI, and the end customer won't notice as it's the Microsoft or Office brand that they buy into. They are not going to advertise what's inside their products with big banners "Powered by OpenAI" in the long run, I think (do they now?)

Re: QLoRA: Efficient Finetuning of Quantized LLMs

#72

Can someone help me understand what quantization means in this context, and why it matters?

GPT-4 ELI5: - 4-bit Quantization: Imagine you have a box of 16 different colored crayons. But you realize that you can draw almost the same picture using only 4 colors. That's what quantization does. It reduces the number of different "colors" (or numbers) that the model uses to represent its knowledge, which saves a lot of space. In this case, they used a special kind of 4-bit quantization, which means they only use…

[deleted]

Re: QLoRA: Efficient Finetuning of Quantized LLMs

#73
post #62
post #8

Earlier quoted context omitted.

Altman’s push for regulatory capture makes so much sense given how fast this field is going. Open models you can run on regular hardware are still behind GPT-4 by some distance but they are closing in at a rate that leads me to believe there’s not much of a moat there. There’s also a ton of promising work on quantization and pruning and other acceleration and compression techniques to make more powerful models run on…

>Open models... are still behind GPT-4 by some distance but they are closing in at a rate that leads me to believe there’s not much of a moat there Commercially, this doesn't matter even if true. As I keep reminding here: moats are rarely about raw tech. Moats are much more about integrations, brand ("Nobody got fired for buying X"), access to API/raw compute... If OpenAI have a superior Office 365 integration they'l…

Open models can talk about "forbidden" topics and can be extended by users. Both of them are significant advantages.

Re: QLoRA: Efficient Finetuning of Quantized LLMs

#74
post #8
post #2

I'm very impressed at the quality of Guanaco 33B, the model that accompanies this paper. You can try it out here: https://huggingface.co/spaces/uwnlp/guanaco-playground-tgi I tried "You are a sentient cheesecake that teaches people SQL, with cheesecake analogies to illustrate different points. Teach me to use count and group by" and got a good result from it: https://twitter.com/simonw/status/1661460336334241794/phot…

Altman’s push for regulatory capture makes so much sense given how fast this field is going. Open models you can run on regular hardware are still behind GPT-4 by some distance but they are closing in at a rate that leads me to believe there’s not much of a moat there. There’s also a ton of promising work on quantization and pruning and other acceleration and compression techniques to make more powerful models run on…

Just to be super clear, do you mean Altman's push for greater regulation, or that he is pushing for actual regulatory capture i.e. corruption of regulating authorities?

Re: QLoRA: Efficient Finetuning of Quantized LLMs

#76
post #2

I'm very impressed at the quality of Guanaco 33B, the model that accompanies this paper. You can try it out here: https://huggingface.co/spaces/uwnlp/guanaco-playground-tgi I tried "You are a sentient cheesecake that teaches people SQL, with cheesecake analogies to illustrate different points. Teach me to use count and group by" and got a good result from it: https://twitter.com/simonw/status/1661460336334241794/phot…

It did fail on: "Burt's father has 3 sons, Jack and John. What's the name of the 3rd son?"

Re: QLoRA: Efficient Finetuning of Quantized LLMs

#77
post #19
post #14

This is off-topic, but are there any communities or congregations (that aren't reddit) based around locally hosted LLMs? I'm asking because while I see a bunch of projects for exposing GGML/LLaMA to OpenAI compatible interfaces, some UIs, etc, I can't really find a good community or resources for the concept in general. I'm working on a front-end for LLMs in general, having re-implemented a working version of OpenAI'…

it's reddit, but /r/LocalLLaMA/

A name that is destined to be obsolete.

Re: QLoRA: Efficient Finetuning of Quantized LLMs

#78
post #33
post #17

Earlier quoted context omitted.

If progress really continues like this any regulation that does not limit open models would be a pointless exercise. It’s only a matter of time before people really crack distributed training algorithms that can be run in a less organized swarm configuration. At that point open trainers could actually train near the frontier of what is possible. Most of the data is open.

> It's only a matter of time before people really crack distributed training algorithms that can be run in a less organized swarm configuration. Why do you think this? This continually failed, and seems extremely unlikely to me. Barring surprising breakthrough, there is inherent communication complexity, and physical limit to communication bandwidth.

We could try something like Civitai is doing already, but automated.

Each node could train the model on a separate concept and then combine the results.

Re: QLoRA: Efficient Finetuning of Quantized LLMs

#79
post #19
post #14

This is off-topic, but are there any communities or congregations (that aren't reddit) based around locally hosted LLMs? I'm asking because while I see a bunch of projects for exposing GGML/LLaMA to OpenAI compatible interfaces, some UIs, etc, I can't really find a good community or resources for the concept in general. I'm working on a front-end for LLMs in general, having re-implemented a working version of OpenAI'…

it's reddit, but /r/LocalLLaMA/

There's also r/oobabooga and r/MachineLearning

Re: QLoRA: Efficient Finetuning of Quantized LLMs

#80
post #73
post #62

Earlier quoted context omitted.

>Open models... are still behind GPT-4 by some distance but they are closing in at a rate that leads me to believe there’s not much of a moat there Commercially, this doesn't matter even if true. As I keep reminding here: moats are rarely about raw tech. Moats are much more about integrations, brand ("Nobody got fired for buying X"), access to API/raw compute... If OpenAI have a superior Office 365 integration they'l…

Open models can talk about "forbidden" topics and can be extended by users. Both of them are significant advantages.

The ability to talk about 'forbidden' topics is also a significant disadvantage. Just wait for the first moral panic targeting open source GPTs. I think that Open Source will exist (barring a legal ban), but the triumphalism is in my mind very unjustified. There's a fair chance closed source will get 99% of this market.
Post reply on HN