Live data from Hacker News

25L Portable NV-linked Dual 3090 LLM Rig

reddit.com

121–126 of 126 posts

Re: 25L Portable NV-linked Dual 3090 LLM Rig

#121
post #37

Earlier quoted context omitted.

I bought a 2nd 3090 2 years ago for like 800eur, still a good price even today I think. It's in my main workstation, and my idea was to always have Ollama running locally. The problem is that once I have a (large-ish) model running, all my VRAM is almost full and GPU struggles to do things like playing back a YouTube video. Lately I haven't used local AI much, also because I stopped using any coding AIs (as they wast…

I run my desktop environment on the iGPU and the AI stuff on the dGPUs.

That's a real good point!

Unfortuatenly, my CPU (5900x) doesn't have an iGPU.

The last 5 years iGPU got a bit out of trend. Now maybe they actually make a lot of sense, as there is a clear use-case which involves having dedicated GPU always in-use which is not gaming (and gaming is different, cause you don't often multi-task while gaming).

I do expect to see a surge in iGPU popularity, or maybe a software improvement to allow having a model always available without constantly hogging the VRAM.

Re: 25L Portable NV-linked Dual 3090 LLM Rig

#122
post #121

Earlier quoted context omitted.

I run my desktop environment on the iGPU and the AI stuff on the dGPUs.

That's a real good point! Unfortuatenly, my CPU (5900x) doesn't have an iGPU. The last 5 years iGPU got a bit out of trend. Now maybe they actually make a lot of sense, as there is a clear use-case which involves having dedicated GPU always in-use which is not gaming (and gaming is different, cause you don't often multi-task while gaming). I do expect to see a surge in iGPU popularity, or maybe a software improvement…

PS: I thought Ollama had a way to use RAM instead of VRAM (?) to keep the model active when not in use, but in my experience that didn't solve the problem.

Re: 25L Portable NV-linked Dual 3090 LLM Rig

#123
post #29

OK, here's my quick critique of the article (having built a similar AM4-based system in 2023 for 2300€): 1) [I thought] The page is blocking cut & paste. Super annoying! 2) The exact mainboard is not specified exactly. There are 4 different boards called "ASUS ROG Strix X670E Gaming" and some of them only have one PCIe x16 slot. None of them can do PCIe x8 when using two GPUs. 3) The shopping link for the mainboard l…

Any reason you wouldn't opt for the 4090 or 5090?

3090 second hand can be found at something like $600.

Re: 25L Portable NV-linked Dual 3090 LLM Rig

#125

Earlier quoted context omitted.

... and this is why napkin calculation is terrible. Even running a GPU at load doesn't mean you are going to use the full wattage. 4 3090 running inference on large model barely uses 350watts combined.

Can you clarify? Even if you down clock the card to 300W, why would running it at load not consume 4x300W?

Inference is often like 200-250w without card clocked down. Then the other cards are like 20w-50w. 4 cards, 1 card is active at once. To get the full 350watt, you need to run parallel inference on the card with multiple users. So if I was using it as a server card and have 10 active users/processes then I might max out the active card. For example, I have a rig with 10 MI50 cards, I believe they are 250w each. Yet I rarely see pass 200w on the active card, they idle at about 20w, so that's 180w + 200w = around 380-400w on full load.

Think of the max watt like a car's max horsepower, a car might make 350HP, it doesn't mean it stays making 350HP all day long, there's a curve to it. At the low end it might be making 170HP and you will need to floor the gas pedal to get to that 350hp. Same with these GPUs. Most people will calculate the gas mileage by finding how much gas a car consumers at it's peak and say, oh, 6mpg when it's making 350hp so with your 20gallon thank, you have a range of 120miles. Which obviously isn't true.

Post reply on HN