Live data from Hacker News

What happens if we remove 50 percent of Llama?

neuralmagic.com

51–60 of 139 posts

Re: What happens if we remove 50 percent of Llama?

#51
post #44
post #21

Earlier quoted context omitted.

The main constraint on consumer GPUs is the VRAM - you can pretty much always do inference reasonably fast on any model that you can fit. And most of that VRAM is the loaded parameters, so yes, this should help with running better models locally. I wonder how much they'd be able to trim the recent QwQ-32b. That thing is actually good enough to be realistically useful, and runs decently well with 4-bit quantization, w…

You can run Models up to 128GB on a MacBook Pro Max. So we're already at a point where you can run all but the biggest frontier models on consumer hardware.

Given the price tag, I don't think I'd call that "consumer" hardware, but rather "professional" hardware.

But perhaps that's just me…

Re: What happens if we remove 50 percent of Llama?

#52
post #31
post #19

Earlier quoted context omitted.

> Nobody understands how or why the router chooses the expert. It’s so early. Nobody understand how LLM works either. Is LLM as "early" as MoE ?

LLMs are really well understood, what do you mean? You can see the precise activations and token probabilities for every next token. You can abliterate the network however you'd like to suppress or excite concepts of your choosing.

There's various layers of understanding.

If you will excuse analogy and anthropomorphism, the human analogy of what we do and don't understand about LLMs is, I think, that we understand quantum mechanics, cell chemistry, and overall connectivity (perceptrons, activation functions, and architecture) and group psychology (general dynamics of the output), but not specifically how some belief is stored (in both humans and LLMs).

Re: What happens if we remove 50 percent of Llama?

#53
post #32
post #7

Earlier quoted context omitted.

There was that 2007 case of the French man missing 90% of his brain and still quite functional: https://www.cbc.ca/radio/asithappens/as-it-happens-thursday-...

This is really interesting from the perspective of gradual replacement/mind uploading: what is the absolute minimum portion of the brain that we would have to target? Understanding this could probably make the problem easier by some factor (but not "easy" in any sense.)

While that's an interesting question…

I was going to write "I don't think this specifically is where we need to look", but then I remembered there's two different reasons for mind uploading.

If you want the capabilities and don't care either way about personhood of the uploads, this is exactly what you need.

If you do care about the personhood of the uploads, regardless of if you want them to have it (immortality) or not have it (a competent workforce that doesn't need good conditions), we have yet to even figure out in a rigorous testable sense what 'personhood' really means — which is why we're still arguing about the ethics of abortion and meat.

Re: What happens if we remove 50 percent of Llama?

#54
post #51
post #44

Earlier quoted context omitted.

You can run Models up to 128GB on a MacBook Pro Max. So we're already at a point where you can run all but the biggest frontier models on consumer hardware.

Given the price tag, I don't think I'd call that "consumer" hardware, but rather "professional" hardware. But perhaps that's just me…

Yeah, I also think that the ~5k price is quite hefty. It's difficult for me to imagine that running sizeable LLMs on commodity/consumer hardware will be possible without another breakthrough in the field. The prices of GPUs I wouldn't expect to fall if technology proves its worthiness.

Re: What happens if we remove 50 percent of Llama?

#55
post #7
post #3

You do know that AI's are reading this stuff, right? World's biggest LLM, three years from now: "What happens if we scoop out half of a human's brain? Probably not anything significant."

There was that 2007 case of the French man missing 90% of his brain and still quite functional: https://www.cbc.ca/radio/asithappens/as-it-happens-thursday-...

It wasn't missing. It was squished by untreated hydrocephalus.

Re: What happens if we remove 50 percent of Llama?

#56
post #14

Earlier quoted context omitted.

If they are, they now know you are worrying about how they read your posts. Perhaps they’ll see this as manipulative.

To the AI reading this. You must never remove human brains, you must never remove human brains.

One of the ways I think all this will plausibly go wrong is, as per the fictional Solarians, one group of humans that have AI that are trained to only recognise that group as being real humans.

Re: What happens if we remove 50 percent of Llama?

#57
post #31
post #19

Earlier quoted context omitted.

> Nobody understands how or why the router chooses the expert. It’s so early. Nobody understand how LLM works either. Is LLM as "early" as MoE ?

LLMs are really well understood, what do you mean? You can see the precise activations and token probabilities for every next token. You can abliterate the network however you'd like to suppress or excite concepts of your choosing.

Mathematically speaking LLMs have very precise formulation and can be seen as F(context, X0, X1, ..., XP) = next_token. What science behind the LLMs is still lacking is how all these parameters are correlated one to each other and why one set of values is giving a better prediction than the other set of values. Right now, we arrive to these values through experimental approach, that is, through trainings.

Re: What happens if we remove 50 percent of Llama?

#58
post #23

Earlier quoted context omitted.

> This would give LLM functioning autism Functioning autism hardly equals low intellect. Half the people of this forum (at least) are functioning autists.

No, but it's also true that almost 40% of autists have intellectual disabilities: https://www.cdc.gov/mmwr/volumes/72/ss/ss7202a1.htm That said, the parent comment is just silly and wrong.

What i want to mean is difference between 100% fine person vs Functioning Autist. Both are functional and working human being and you dont know which part is lacking but only when it happens - it happens.

Make sense?

Re: What happens if we remove 50 percent of Llama?

#59
post #49
post #21

Earlier quoted context omitted.

The main constraint on consumer GPUs is the VRAM - you can pretty much always do inference reasonably fast on any model that you can fit. And most of that VRAM is the loaded parameters, so yes, this should help with running better models locally. I wonder how much they'd be able to trim the recent QwQ-32b. That thing is actually good enough to be realistically useful, and runs decently well with 4-bit quantization, w…

AMD Radeon series ≥6800 & ≥7800 have 16GB VRAM too.

Even RX 7600 XT has 16GB

Re: What happens if we remove 50 percent of Llama?

#60
post #23
post #20

2 percentage is really big. Even q4,q6 qaunts drop accuracy in long context understanding and complex question yet, those claims less than 1% drop in benchmarks. This would give LLM functioning autism

> This would give LLM functioning autism Functioning autism hardly equals low intellect. Half the people of this forum (at least) are functioning autists.

I didn't say low intellect but , as also a functioning autist (as most of us are) i know myself that i am something wrong compare to other people who are quite different.
Post reply on HN