Earlier quoted context omitted.
I think the QLoRA paper https://arxiv.org/pdf/2305.14314 paper also showed LoRA on all MLP + Attention layers > all MLP layers > just Attention layers. Other papers show finetuning a select few layers can also work well.
Any real world performance comparison between QLoRa and LoRa?
LoRA Learns Less and Forgets Less
51–60 of 65 posts
Re: LoRA Learns Less and Forgets Less
#52I really wish people would be more careful about choosing names for these things. LoRa has been a popular wireless protocol for like 10 years.
Re: LoRA Learns Less and Forgets Less
#53Earlier quoted context omitted.
...but they should know how to use search engines...
at some point we are going to run out of easily pronounceable abbreviations that are unique. Perhaps that point is actually in the past and we should just acknowledge it and move on. Although I guess it could have been Lorall - oops, that's a character in World of Warcraft.
Re: LoRA Learns Less and Forgets Less
#54Re: LoRA Learns Less and Forgets Less
#55This paper has 12 authors, which fascinates me to no end for some unexplainable reason. How does it work? Is it a common occurrence to have this many people working on a submission? Did each of them get at least a paragraph in edgewise?
The general criteria for authorship require including the people who worked on the experiments and data for the paper, which can be more important contribution than most of the text in that paper. In other experimental fields, there are papers with dozens or even hundreds of authors, because it can take many people to get to a measurement of a single number in the paper.
Re: LoRA Learns Less and Forgets Less
#56I really wish people would be more careful about choosing names for these things. LoRa has been a popular wireless protocol for like 10 years.
Re: LoRA Learns Less and Forgets Less
#57Earlier quoted context omitted.
Yes, but this is LoRA, clearly not LoRa.
PP has a point though. I entered "LoRA" on Google, DuckDuckGo, Startpage and Bing, and all returned results in all first pages were about the communication protocol (1). They could have inferred my interests from previous searches, but I never used Bing in the last year or so, so it seems to me someone didn't care about name clashes. (1) well, except Google which -surprise- returned about mid page an ad of a local qu…
Re: LoRA Learns Less and Forgets Less
#58I really wish people would be more careful about choosing names for these things. LoRa has been a popular wireless protocol for like 10 years.
Re: LoRA Learns Less and Forgets Less
#59This paper has 12 authors, which fascinates me to no end for some unexplainable reason. How does it work? Is it a common occurrence to have this many people working on a submission? Did each of them get at least a paragraph in edgewise?
I raise you the Gemini paper https://arxiv.org/abs/2312.11805
Re: LoRA Learns Less and Forgets Less
#60Earlier quoted context omitted.
99% of the engineers who are still working in ML (Machine Language) would. A much smaller percent among those who write in ML (the functional programming language) probably, though.
Even if the ML engineers know about the wireless protocol (and I doubt that many do), the scientists/researchers who develop these models probably don't. They are completely different domain. The lead author on this paper is basically a neuroscientist; some of the other are technically computer scientists, but probably have little hands-on experience with networking beyond whatever they did in undergrad.