QLoRA: Efficient Finetuning of Quantized LLMs
1–10 of 116 posts
Re: QLoRA: Efficient Finetuning of Quantized LLMs
#2You can try it out here: https://huggingface.co/spaces/uwnlp/guanaco-playground-tgi
I tried "You are a sentient cheesecake that teaches people SQL, with cheesecake analogies to illustrate different points. Teach me to use count and group by" and got a good result from it: https://twitter.com/simonw/status/1661460336334241794/photo/...
Re: QLoRA: Efficient Finetuning of Quantized LLMs
#3I'm very impressed at the quality of Guanaco 33B, the model that accompanies this paper. You can try it out here: https://huggingface.co/spaces/uwnlp/guanaco-playground-tgi I tried "You are a sentient cheesecake that teaches people SQL, with cheesecake analogies to illustrate different points. Teach me to use count and group by" and got a good result from it: https://twitter.com/simonw/status/1661460336334241794/phot…
Excerpt: “As a sentient cow with a PhD in moo-matics, I am happy to explain why 2+2 equals 4, my dear hooman friend… In moo-matical terms, each number is actually made up of smaller units called digits.”
I approve.
Re: QLoRA: Efficient Finetuning of Quantized LLMs
#4Re: QLoRA: Efficient Finetuning of Quantized LLMs
#5Re: QLoRA: Efficient Finetuning of Quantized LLMs
#6Is lemon-picked a real phrase or did they use GPT to generate the abstract? The term is “cherry-picked”.
Re: QLoRA: Efficient Finetuning of Quantized LLMs
#7Re: QLoRA: Efficient Finetuning of Quantized LLMs
#8I'm very impressed at the quality of Guanaco 33B, the model that accompanies this paper. You can try it out here: https://huggingface.co/spaces/uwnlp/guanaco-playground-tgi I tried "You are a sentient cheesecake that teaches people SQL, with cheesecake analogies to illustrate different points. Teach me to use count and group by" and got a good result from it: https://twitter.com/simonw/status/1661460336334241794/phot…
There’s also a ton of promising work on quantization and pruning and other acceleration and compression techniques to make more powerful models run on smaller devices. So far the focus has been on just getting these things to work, not efficiency. There’s probably a lot of fruit to be picked here.
A few more years and a gaming PC may be at GPT-4 level or maybe even better.
No everyone won’t run their own models but it shows that there will end up being many commercial apps and services and they won’t all have to use OpenAI’s API. There’s going to be lots of competition. Unless of course it’s regulated away.
Re: QLoRA: Efficient Finetuning of Quantized LLMs
#9Over 1,000 models finetuned! Finetuning 65B models on consumer hardware in under a day, with full 16bit finetune performance.
4bit does it again!
Re: QLoRA: Efficient Finetuning of Quantized LLMs
#10Is lemon-picked a real phrase or did they use GPT to generate the abstract? The term is “cherry-picked”.