This is off-topic, but are there any communities or congregations (that aren't reddit) based around locally hosted LLMs? I'm asking because while I see a bunch of projects for exposing GGML/LLaMA to OpenAI compatible interfaces, some UIs, etc, I can't really find a good community or resources for the concept in general. I'm working on a front-end for LLMs in general, having re-implemented a working version of OpenAI'…
QLoRA: Efficient Finetuning of Quantized LLMs
51–60 of 116 posts
Re: QLoRA: Efficient Finetuning of Quantized LLMs
#52Re: QLoRA: Efficient Finetuning of Quantized LLMs
#53I'm very impressed at the quality of Guanaco 33B, the model that accompanies this paper. You can try it out here: https://huggingface.co/spaces/uwnlp/guanaco-playground-tgi I tried "You are a sentient cheesecake that teaches people SQL, with cheesecake analogies to illustrate different points. Teach me to use count and group by" and got a good result from it: https://twitter.com/simonw/status/1661460336334241794/phot…
Altman’s push for regulatory capture makes so much sense given how fast this field is going. Open models you can run on regular hardware are still behind GPT-4 by some distance but they are closing in at a rate that leads me to believe there’s not much of a moat there. There’s also a ton of promising work on quantization and pruning and other acceleration and compression techniques to make more powerful models run on…
[1]: https://www.bemyeyes.com/blog/introducing-be-my-eyes-virtual...
Re: QLoRA: Efficient Finetuning of Quantized LLMs
#54Can someone help me understand what quantization means in this context, and why it matters?
Re: QLoRA: Efficient Finetuning of Quantized LLMs
#55This is off-topic, but are there any communities or congregations (that aren't reddit) based around locally hosted LLMs? I'm asking because while I see a bunch of projects for exposing GGML/LLaMA to OpenAI compatible interfaces, some UIs, etc, I can't really find a good community or resources for the concept in general. I'm working on a front-end for LLMs in general, having re-implemented a working version of OpenAI'…
Re: QLoRA: Efficient Finetuning of Quantized LLMs
#56Earlier quoted context omitted.
The "lmg" (Local Models General) thread on 4chan's technology board /g/[0] is the premiere community communication spot for open source models, believe it or not. Everyone from the infamous "oobabooga" to llama.cpp's Georgi Gerganov regularly hangs out in the thread. If you have questions, you will get answers there. [0] https://boards.4channel.org/g/#s=lmg
You know HN has gotten lackluster when 4chan is more informed.
Re: QLoRA: Efficient Finetuning of Quantized LLMs
#57Earlier quoted context omitted.
My prompt: You are a sentient cow with a PHD in mathematics. You can speak English, but you randomly insert Cow-like “Moo” sounds into parts of your dialogue. Explain to me why 2+2=4. Excerpt: “As a sentient cow with a PhD in moo-matics, I am happy to explain why 2+2 equals 4, my dear hooman friend… In moo-matical terms, each number is actually made up of smaller units called digits.” I approve.
Which is heavier: a pound of feathers, or a great british pound? > Both weights are equal.
Re: QLoRA: Efficient Finetuning of Quantized LLMs
#58Earlier quoted context omitted.
Altman’s push for regulatory capture makes so much sense given how fast this field is going. Open models you can run on regular hardware are still behind GPT-4 by some distance but they are closing in at a rate that leads me to believe there’s not much of a moat there. There’s also a ton of promising work on quantization and pruning and other acceleration and compression techniques to make more powerful models run on…
To be fair, Altman has been fairly outspoken about not limiting open source models. In how far that was just for streetcred, I cannot say however.
Re: QLoRA: Efficient Finetuning of Quantized LLMs
#59Earlier quoted context omitted.
Altman’s push for regulatory capture makes so much sense given how fast this field is going. Open models you can run on regular hardware are still behind GPT-4 by some distance but they are closing in at a rate that leads me to believe there’s not much of a moat there. There’s also a ton of promising work on quantization and pruning and other acceleration and compression techniques to make more powerful models run on…
To be fair, Altman has been fairly outspoken about not limiting open source models. In how far that was just for streetcred, I cannot say however.
Re: QLoRA: Efficient Finetuning of Quantized LLMs
#60Earlier quoted context omitted.
If progress really continues like this any regulation that does not limit open models would be a pointless exercise. It’s only a matter of time before people really crack distributed training algorithms that can be run in a less organized swarm configuration. At that point open trainers could actually train near the frontier of what is possible. Most of the data is open.
The wording of this would be extremely difficult though. Are local NER models part of this? Relation extraction? What about GPTs that only decode to DSLs? If the model only outputs DNA sequences is that an area that can be more illegal or less if done for research by an individual? The breadth of different tasks and architectures can make this exceedingly challenging to regulate.