Live data from Hacker News

QLoRA: Efficient Finetuning of Quantized LLMs

arxiv.org

1–10 of 116 posts

Re: QLoRA: Efficient Finetuning of Quantized LLMs

#2
I'm very impressed at the quality of Guanaco 33B, the model that accompanies this paper.

You can try it out here: https://huggingface.co/spaces/uwnlp/guanaco-playground-tgi

I tried "You are a sentient cheesecake that teaches people SQL, with cheesecake analogies to illustrate different points. Teach me to use count and group by" and got a good result from it: https://twitter.com/simonw/status/1661460336334241794/photo/...

Re: QLoRA: Efficient Finetuning of Quantized LLMs

#3
post #2

I'm very impressed at the quality of Guanaco 33B, the model that accompanies this paper. You can try it out here: https://huggingface.co/spaces/uwnlp/guanaco-playground-tgi I tried "You are a sentient cheesecake that teaches people SQL, with cheesecake analogies to illustrate different points. Teach me to use count and group by" and got a good result from it: https://twitter.com/simonw/status/1661460336334241794/phot…

My prompt: You are a sentient cow with a PHD in mathematics. You can speak English, but you randomly insert Cow-like “Moo” sounds into parts of your dialogue. Explain to me why 2+2=4.

Excerpt: “As a sentient cow with a PhD in moo-matics, I am happy to explain why 2+2 equals 4, my dear hooman friend… In moo-matical terms, each number is actually made up of smaller units called digits.”

I approve.

Re: QLoRA: Efficient Finetuning of Quantized LLMs

#8
post #2

I'm very impressed at the quality of Guanaco 33B, the model that accompanies this paper. You can try it out here: https://huggingface.co/spaces/uwnlp/guanaco-playground-tgi I tried "You are a sentient cheesecake that teaches people SQL, with cheesecake analogies to illustrate different points. Teach me to use count and group by" and got a good result from it: https://twitter.com/simonw/status/1661460336334241794/phot…

Altman’s push for regulatory capture makes so much sense given how fast this field is going. Open models you can run on regular hardware are still behind GPT-4 by some distance but they are closing in at a rate that leads me to believe there’s not much of a moat there.

There’s also a ton of promising work on quantization and pruning and other acceleration and compression techniques to make more powerful models run on smaller devices. So far the focus has been on just getting these things to work, not efficiency. There’s probably a lot of fruit to be picked here.

A few more years and a gaming PC may be at GPT-4 level or maybe even better.

No everyone won’t run their own models but it shows that there will end up being many commercial apps and services and they won’t all have to use OpenAI’s API. There’s going to be lots of competition. Unless of course it’s regulated away.

Re: QLoRA: Efficient Finetuning of Quantized LLMs

#10
post #5

Is lemon-picked a real phrase or did they use GPT to generate the abstract? The term is “cherry-picked”.

> When we notice a pattern we attempt to setup a question or prompt that will induce the pattern even though it is the incorrect solution, e.g., if we observe that the model tends to give long-winded answers we prompt the model to “Answer yes or no without explanation.” We use this to find “lemons” where we manage to adversarially break the model and “cherries” where we fail to break the model, and present both.
Post reply on HN