Live data from Hacker News

QLoRA: Efficient Finetuning of Quantized LLMs

arxiv.org

51–60 of 116 posts

Re: QLoRA: Efficient Finetuning of Quantized LLMs

#51
post #14

This is off-topic, but are there any communities or congregations (that aren't reddit) based around locally hosted LLMs? I'm asking because while I see a bunch of projects for exposing GGML/LLaMA to OpenAI compatible interfaces, some UIs, etc, I can't really find a good community or resources for the concept in general. I'm working on a front-end for LLMs in general, having re-implemented a working version of OpenAI'…

You might want to try some of the discord channels connected to some of the repos. i.e. GPT4All https://github.com/nomic-ai/gpt4all scroll down for the discord link.

Re: QLoRA: Efficient Finetuning of Quantized LLMs

#53
post #8
post #2

I'm very impressed at the quality of Guanaco 33B, the model that accompanies this paper. You can try it out here: https://huggingface.co/spaces/uwnlp/guanaco-playground-tgi I tried "You are a sentient cheesecake that teaches people SQL, with cheesecake analogies to illustrate different points. Teach me to use count and group by" and got a good result from it: https://twitter.com/simonw/status/1661460336334241794/phot…

Altman’s push for regulatory capture makes so much sense given how fast this field is going. Open models you can run on regular hardware are still behind GPT-4 by some distance but they are closing in at a rate that leads me to believe there’s not much of a moat there. There’s also a ton of promising work on quantization and pruning and other acceleration and compression techniques to make more powerful models run on…

I wonder if there's a campaign that Americans can get behind to counter Altman's push for regulation. Such a thing might get me to engage in politics again. I don't want the current pace of innovation in running these models on one's own computer to slow down, because it could be a good thing for assistive technology, e.g. running something like the Be My Eyes Virtual Volunteer [1] on one's own computer.

[1]: https://www.bemyeyes.com/blog/introducing-be-my-eyes-virtual...

Re: QLoRA: Efficient Finetuning of Quantized LLMs

#55
post #14

This is off-topic, but are there any communities or congregations (that aren't reddit) based around locally hosted LLMs? I'm asking because while I see a bunch of projects for exposing GGML/LLaMA to OpenAI compatible interfaces, some UIs, etc, I can't really find a good community or resources for the concept in general. I'm working on a front-end for LLMs in general, having re-implemented a working version of OpenAI'…

[deleted]

Re: QLoRA: Efficient Finetuning of Quantized LLMs

#56

Earlier quoted context omitted.

The "lmg" (Local Models General) thread on 4chan's technology board /g/[0] is the premiere community communication spot for open source models, believe it or not. Everyone from the infamous "oobabooga" to llama.cpp's Georgi Gerganov regularly hangs out in the thread. If you have questions, you will get answers there. [0] https://boards.4channel.org/g/#s=lmg

You know HN has gotten lackluster when 4chan is more informed.

HN is hardly the right place for an ongoing discussion

Re: QLoRA: Efficient Finetuning of Quantized LLMs

#57
post #39
post #3

Earlier quoted context omitted.

My prompt: You are a sentient cow with a PHD in mathematics. You can speak English, but you randomly insert Cow-like “Moo” sounds into parts of your dialogue. Explain to me why 2+2=4. Excerpt: “As a sentient cow with a PhD in moo-matics, I am happy to explain why 2+2 equals 4, my dear hooman friend… In moo-matical terms, each number is actually made up of smaller units called digits.” I approve.

Which is heavier: a pound of feathers, or a great british pound? > Both weights are equal.

So same answer as free ChatGPT, only terser.

Re: QLoRA: Efficient Finetuning of Quantized LLMs

#58
post #11
post #8

Earlier quoted context omitted.

Altman’s push for regulatory capture makes so much sense given how fast this field is going. Open models you can run on regular hardware are still behind GPT-4 by some distance but they are closing in at a rate that leads me to believe there’s not much of a moat there. There’s also a ton of promising work on quantization and pruning and other acceleration and compression techniques to make more powerful models run on…

To be fair, Altman has been fairly outspoken about not limiting open source models. In how far that was just for streetcred, I cannot say however.

Open models means that OpenAI has access to and can learn from them. If the move is to target commercial competition, then this doesn't preclude it.

Re: QLoRA: Efficient Finetuning of Quantized LLMs

#59
post #11
post #8

Earlier quoted context omitted.

Altman’s push for regulatory capture makes so much sense given how fast this field is going. Open models you can run on regular hardware are still behind GPT-4 by some distance but they are closing in at a rate that leads me to believe there’s not much of a moat there. There’s also a ton of promising work on quantization and pruning and other acceleration and compression techniques to make more powerful models run on…

To be fair, Altman has been fairly outspoken about not limiting open source models. In how far that was just for streetcred, I cannot say however.

I won't take him seriously on this - if one really believes in the case for regulation, than there's no good reason to exclude open source models. Let's take the most pessimistic possibility, that they have a slower rate of progress than commercial models, and real progress depends on hardware. We still end up with the same result - every 'risky' point commercial models get to, open source will get to as well.

Re: QLoRA: Efficient Finetuning of Quantized LLMs

#60
post #49
post #17

Earlier quoted context omitted.

If progress really continues like this any regulation that does not limit open models would be a pointless exercise. It’s only a matter of time before people really crack distributed training algorithms that can be run in a less organized swarm configuration. At that point open trainers could actually train near the frontier of what is possible. Most of the data is open.

The wording of this would be extremely difficult though. Are local NER models part of this? Relation extraction? What about GPTs that only decode to DSLs? If the model only outputs DNA sequences is that an area that can be more illegal or less if done for research by an individual? The breadth of different tasks and architectures can make this exceedingly challenging to regulate.

[deleted]
Post reply on HN