Live data from Hacker News

25L Portable NV-linked Dual 3090 LLM Rig

reddit.com

91–100 of 126 posts

Re: 25L Portable NV-linked Dual 3090 LLM Rig

#91
I was going to say you need an extension cable. My first dual 3090 build I had three issues. First was the pcie extension wouldn't support gen4, so I had to change to gen3 in the bios. Second issue was that depending on which slot, you couldn't get x16/x16 and it would drop to x16/x8 unless you had it configured right. Third, I finally gave up and just had the card resting first inside the case and then outside which if fan kicks up, it'll jiggle around, so I had to make some makeshift holder to keep the card sitting there.

Re: 25L Portable NV-linked Dual 3090 LLM Rig

#92
post #29

OK, here's my quick critique of the article (having built a similar AM4-based system in 2023 for 2300€): 1) [I thought] The page is blocking cut & paste. Super annoying! 2) The exact mainboard is not specified exactly. There are 4 different boards called "ASUS ROG Strix X670E Gaming" and some of them only have one PCIe x16 slot. None of them can do PCIe x8 when using two GPUs. 3) The shopping link for the mainboard l…

> 3) The shopping link for the mainboard leads to the "ASUS ROG Strix X670E-E Gaming" model. This model can use the 2nd PCIe 5.0 port at only x4 speeds. The RTX 3090 can only do PCIe 4.0 of course so it will run at PCIe 4.0 x4. If you choose a desktop mainboard for having two GPUs, make sure it can run at PCIe x8 speeds when using both GPU slots! Having NVLink between the GPUs is not a replacement for having a fast connection between the CPU+RAM and the GPU and its VRAM.

Forgive a noob question: I thought the connection to the GPU was actually fairly unimportant once the model was loaded, because sending input to the model and getting a response is low bandwidth? So it might matter if you're changing models a lot or doing a model that can work on video, but otherwise I thought it didn't really matter.

Re: 25L Portable NV-linked Dual 3090 LLM Rig

#93

Earlier quoted context omitted.

> The page is blocking cut & paste. Super annoying! I've been running Don't F* With Paste* for years for this https://chromewebstore.google.com/detail/dont-f-with-paste/n...

Interesting. I guess our content-based marketing pages need to move to canvas-based rendering. That's probably bum too. Straight to serving up jpgs.

thankfully most web browsing will be done by LLMs soon and that won't stop them, good riddance to the mess of a web that google has created

Re: 25L Portable NV-linked Dual 3090 LLM Rig

#94
post #48
post #29

OK, here's my quick critique of the article (having built a similar AM4-based system in 2023 for 2300€): 1) [I thought] The page is blocking cut & paste. Super annoying! 2) The exact mainboard is not specified exactly. There are 4 different boards called "ASUS ROG Strix X670E Gaming" and some of them only have one PCIe x16 slot. None of them can do PCIe x8 when using two GPUs. 3) The shopping link for the mainboard l…

Yeah, this page seems to be not great for beginners and also useless for people with experience. A 2x 3090 build is okay for inference, but even with nvlink you're a bit handicapped for training. You're much better off with getting a 4090 48GB from China for $2.5k and just using that. Example: https://www.alibaba.com/trade/search?keywords=4090+48gb&pric... Also, this phrasing is concerning: > WARNING - these componen…

What an indictment on NVidia market segmentation that there's an industry doing aftermarket VRAM upgrades on gaming cards due their intentionally hobbled VRAM.

I wish AMD and Intel Arc would step up their game.

Re: 25L Portable NV-linked Dual 3090 LLM Rig

#95
post #81

There was an interesting post to r/LocalLLaMA yesterday from someone running inference mostly on CPU: https://carteakey.dev/optimizing%20gpt-oss-120b-local%20infe... One of the observations is how much difference memory speed and bandwidth makes, even for CPU inference. Obviously a CPU isn't going to match a GPU for inference speed, but it's an affordable way to run much larger models than you can fit in 24GB or even…

Outside of prompt processing, the only reason GPU's are better than CPU's for inference is memory bandwidth, the performance of apple M* devices at inference is a consequence of this, not of their UMA.

Re: 25L Portable NV-linked Dual 3090 LLM Rig

#96
post #13

Earlier quoted context omitted.

I am curious about the setup of 14 GPUs - what kind of platform (motherboard) do you use to support so many PCIe lanes? And do you even have a chassis? Is it rack-mounted? Thanks!

I used a large supermicro server chassis, a dual Xeon motherboard with 7 8 lane PCI Express slots, all the ram it would take (bought second hand), splitters, four massive powersupplies. I extended the server chassis with aluminum angle riveted onto the base. It could be rack mounted but I'd hate to be the person lifting it in. The 3090s were a mix, 10 of the same type (small, and with blower style fans on them) and 4…

Thanks that is very inspiring. I thought there are no blower type consumer GPUs, but apparently they exist!

Re: 25L Portable NV-linked Dual 3090 LLM Rig

#97
post #87

Earlier quoted context omitted.

I have rig of 7 3090s that I bought from crypto bros, they are lasting quite alright and have been chugging along fine for the last 2 years. GPUs are electronic devices not mechanical devices, they rarely blow up.

How do you have a rig that fits that many cards?? those things take 3 slots apiece. Pictures, or it never happened! :D

you get a motherboard designed for the purpose (many pcie slots) and a case (usually open frame) that holds that many cards. riser cables are used so every card doesnt plug directly into the motherboard

Re: 25L Portable NV-linked Dual 3090 LLM Rig

#98
post #48

Earlier quoted context omitted.

Yeah, this page seems to be not great for beginners and also useless for people with experience. A 2x 3090 build is okay for inference, but even with nvlink you're a bit handicapped for training. You're much better off with getting a 4090 48GB from China for $2.5k and just using that. Example: https://www.alibaba.com/trade/search?keywords=4090+48gb&pric... Also, this phrasing is concerning: > WARNING - these componen…

What an indictment on NVidia market segmentation that there's an industry doing aftermarket VRAM upgrades on gaming cards due their intentionally hobbled VRAM. I wish AMD and Intel Arc would step up their game.

Intel Arc Pro B60 will come in a 48GB dual-GPU model. So yeah, hardware is gonna be there, and the 24GB model will be $599 from Sparkle. I assume 48GB will be cheaper than a hacked RTX 4090.

Look at this: https://www.maxsun.com/products/intel-arc-pro-b60-dual-48g-t... https://www.sparkle.com.tw/files/20250618145718157.pdf

Re: 25L Portable NV-linked Dual 3090 LLM Rig

#99

Earlier quoted context omitted.

On my 2x 3090s I am running glm4.5 air q1 and it runs at ~300pp and 20/30 tk/s works pretty well with roo code on vscode, rarely misses tool calls and produces decent quality code. I also tried to use it with claude code with claude code router and it's pretty fast. Roo code uses bigger contexts, so it's quite slower than claude code in general, but I like the workflow better. this is my snippet for llama-swap ``` mo…

What is llama-swap? Been looking for more details about software configs on https://llamabuilds.ai

https://github.com/mostlygeek/llama-swap

it's a transparent proxy that automatically launches your selected model with your preferred inference server so that you don't need to manually start/stop the server when you want to switch model

so, let's say I have configured roo code to use qwen3 30ba3b as the orchestrator and glm4.5 air as coder, roo code would call the proxy server with model "qwen3" when using orchestrator mode and then kill llama.cpp with qwen3 and restart it with "glm4.5air"

Re: 25L Portable NV-linked Dual 3090 LLM Rig

#100
post #14

Earlier quoted context omitted.

I am exploring options just for fun. a used 3090 is around $900 on ebay. a used rtx 6000 ADA is around $5k 4 3090s are slower at inference and worse at training than 1 rtx 6000. 4x3090 would consume 1400W at load. Rtx 6000 would consume 300W at load. If you god forbid live in California and your power averages 45 cents per kwh, 4x3090 would be $1500+ more per year to operate than a single RTX 6000[0] [0] Back of the…

... and this is why napkin calculation is terrible. Even running a GPU at load doesn't mean you are going to use the full wattage. 4 3090 running inference on large model barely uses 350watts combined.

Can you clarify? Even if you down clock the card to 300W, why would running it at load not consume 4x300W?
Post reply on HN