Live data from Hacker News

LoRA Learns Less and Forgets Less

arxiv.org

51–60 of 65 posts

Re: LoRA Learns Less and Forgets Less

#51
post #40

Earlier quoted context omitted.

I think the QLoRA paper https://arxiv.org/pdf/2305.14314 paper also showed LoRA on all MLP + Attention layers > all MLP layers > just Attention layers. Other papers show finetuning a select few layers can also work well.

Any real world performance comparison between QLoRa and LoRa?

The QLoRA paper itself provided some cool benchmarks across many many experiments - QLoRA is near equivalent to LoRA, with it sometimes exceeding or losing 1-2% accuracy (it depends on the use case)

Re: LoRA Learns Less and Forgets Less

#53

Earlier quoted context omitted.

...but they should know how to use search engines...

at some point we are going to run out of easily pronounceable abbreviations that are unique. Perhaps that point is actually in the past and we should just acknowledge it and move on. Although I guess it could have been Lorall - oops, that's a character in World of Warcraft.

Old concepts become obsolete anyway. People can start reusing VCR, etc.

Re: LoRA Learns Less and Forgets Less

#54
post #27

Earlier quoted context omitted.

Not sure about "popular" 99% of ML engineers wouldn't know what it is.

...but they should know how to use search engines...

As should have the IoT people to not conflict with the decades old LORA name used for Level of Repair Analysis.

Re: LoRA Learns Less and Forgets Less

#55
post #33

This paper has 12 authors, which fascinates me to no end for some unexplainable reason. How does it work? Is it a common occurrence to have this many people working on a submission? Did each of them get at least a paragraph in edgewise?

The general criteria for authorship require including the people who worked on the experiments and data for the paper, which can be more important contribution than most of the text in that paper. In other experimental fields, there are papers with dozens or even hundreds of authors, because it can take many people to get to a measurement of a single number in the paper.

Thanks, this is the bit I've been missing.

Re: LoRA Learns Less and Forgets Less

#56

I really wish people would be more careful about choosing names for these things. LoRa has been a popular wireless protocol for like 10 years.

The scale of human knowledge, and the pace we're increasing it at, means we probably can't have unique names for everything any more if we aim for short, pronounceable things. "Lora" must have dozens of meanings across different domains. The fact you recognize it from another one is just something you have to deal with.

Re: LoRA Learns Less and Forgets Less

#57
post #11

Earlier quoted context omitted.

Yes, but this is LoRA, clearly not LoRa.

PP has a point though. I entered "LoRA" on Google, DuckDuckGo, Startpage and Bing, and all returned results in all first pages were about the communication protocol (1). They could have inferred my interests from previous searches, but I never used Bing in the last year or so, so it seems to me someone didn't care about name clashes. (1) well, except Google which -surprise- returned about mid page an ad of a local qu…

Most of the time you have to add a keyword as to what it is related. We cannot expect everything to have unique names, unless we are perfectly fine with random pronounceable strings as names.

Re: LoRA Learns Less and Forgets Less

#58

I really wish people would be more careful about choosing names for these things. LoRa has been a popular wireless protocol for like 10 years.

The other side of this is that if you become paralyzed by decisions because 0.001% of people are bothered by it, you're not gonna make it.

Re: LoRA Learns Less and Forgets Less

#59
post #35
post #33

This paper has 12 authors, which fascinates me to no end for some unexplainable reason. How does it work? Is it a common occurrence to have this many people working on a submission? Did each of them get at least a paragraph in edgewise?

I raise you the Gemini paper https://arxiv.org/abs/2312.11805

Goes to show how much money is being poured into this stuff.

Re: LoRA Learns Less and Forgets Less

#60

Earlier quoted context omitted.

99% of the engineers who are still working in ML (Machine Language) would. A much smaller percent among those who write in ML (the functional programming language) probably, though.

Even if the ML engineers know about the wireless protocol (and I doubt that many do), the scientists/researchers who develop these models probably don't. They are completely different domain. The lead author on this paper is basically a neuroscientist; some of the other are technically computer scientists, but probably have little hands-on experience with networking beyond whatever they did in undergrad.

[deleted]
Post reply on HN