Live data from Hacker News

25L Portable NV-linked Dual 3090 LLM Rig

reddit.com

51–60 of 126 posts

Re: 25L Portable NV-linked Dual 3090 LLM Rig

#51
post #50

I built a similar system, meanwhile I've sold one of the RTX 3090's. Local inference is fun and feels liberating, but it's also slow, and once I was used to the immense power of the giant hosted models, the fun quickly disappeared. I've kept a single GPU to still be able to play a bit with light local models, but not anymore for serious use.

I have a similar setup as the author with 2x 3090s. The issue is not that it's slow. 20-30 tk/s is perfectly acceptable to me. The issue is that the quality of the models that I'm able to self-host pales in comparison to that of SOTA hosted models. They hallucinate more, don't follow prompts as well, and simply generate overall worse quality content. These are issues that plague all "AI" models, but they are particul…

> 20-30 tk/s

or ~2.2M tk/day. This is how we should be thinking about it imho.

Re: 25L Portable NV-linked Dual 3090 LLM Rig

#52
post #29

OK, here's my quick critique of the article (having built a similar AM4-based system in 2023 for 2300€): 1) [I thought] The page is blocking cut & paste. Super annoying! 2) The exact mainboard is not specified exactly. There are 4 different boards called "ASUS ROG Strix X670E Gaming" and some of them only have one PCIe x16 slot. None of them can do PCIe x8 when using two GPUs. 3) The shopping link for the mainboard l…

Sorry for going off topic. But your insight will be helpful on my build

I'm thinking about a low budget system, which will be using

1.X99 D8 MAX LGA2011-3 Motherboard - It has 4 pcie 3.0 x16 slots, dual cpu socket. They are priced around $260 with both the cpu

2. 4X AMD MI50 32G cards - They are old now, but they have 32 gigs of vram and also can be sources at $110 each

The whole setup would not cost more than $1000, is it a right build ? or something more performant can be built within this budget ?

Re: 25L Portable NV-linked Dual 3090 LLM Rig

#53

I built a similar system, meanwhile I've sold one of the RTX 3090's. Local inference is fun and feels liberating, but it's also slow, and once I was used to the immense power of the giant hosted models, the fun quickly disappeared. I've kept a single GPU to still be able to play a bit with light local models, but not anymore for serious use.

If you have a 24 gb 3090. Try out qwen:30b-a3b-instruct-2507-q4_K_M ( ollama )

It's pretty good.

Re: 25L Portable NV-linked Dual 3090 LLM Rig

#54
post #35

Earlier quoted context omitted.

>I am exploring options just for fun. Since you're exploring options just for fun, out of curiosity, would you rent it out whenever you're not using it yourself, so it's not just sitting idle? (Could be noisy and loud). You'd be able to use your computer for other work at the same time and stop whenever you wanted to use it yourself.

It depends. At my electricity cost, 1 hour of 3090 or 1 hour of Rtx 6000 would cost the same 0.45 Just checked vast.ai. I will be losing money with 3090 at my electricity cost and making a tiny bit with rtx 6000. Like with boats it’s probably better to rent GPUs then buy them

(you should also be compensated for the noise and inconvenience from it, not only electricity.) It sounds like you might rent it out if the rental price were higher.

Re: 25L Portable NV-linked Dual 3090 LLM Rig

#55
post #50

I built a similar system, meanwhile I've sold one of the RTX 3090's. Local inference is fun and feels liberating, but it's also slow, and once I was used to the immense power of the giant hosted models, the fun quickly disappeared. I've kept a single GPU to still be able to play a bit with light local models, but not anymore for serious use.

I have a similar setup as the author with 2x 3090s. The issue is not that it's slow. 20-30 tk/s is perfectly acceptable to me. The issue is that the quality of the models that I'm able to self-host pales in comparison to that of SOTA hosted models. They hallucinate more, don't follow prompts as well, and simply generate overall worse quality content. These are issues that plague all "AI" models, but they are particul…

On my 2x 3090s I am running glm4.5 air q1 and it runs at ~300pp and 20/30 tk/s works pretty well with roo code on vscode, rarely misses tool calls and produces decent quality code.

I also tried to use it with claude code with claude code router and it's pretty fast. Roo code uses bigger contexts, so it's quite slower than claude code in general, but I like the workflow better.

this is my snippet for llama-swap

``` models: "glm45-air": healthCheckTimeout: 300 cmd: | llama.cpp/build/bin/llama-server -hf unsloth/GLM-4.5-Air-GGUF:IQ1_M --split-mode layer --tensor-split 0.48,0.52 --flash-attn on -c 82000 --ubatch-size 512 --cache-type-k q4_1 --cache-type-v q4_1 -ngl 99 --threads -1 --port ${PORT} --host 0.0.0.0 --no-mmap -hfd mradermacher/GLM-4.5-DRAFT-0.6B-v3.0-i1-GGUF:Q6_K -ngld 99 --kv-unified ```

Re: 25L Portable NV-linked Dual 3090 LLM Rig

#56
post #14

I'm a huge fan of OpenRouter and their interface for solid LLM's but I recently jumped into fine tuning / modifying my own vision models for FPV drone detection (just for fun) and my daily workstation and it's 2080 just wasn't good enough. Even in 2025 it's cool how solid a setup dual 3090's still are. nvlink is an absolute must but it's incredibly powerful. I'm able to run the latest Mistral thinking models and rela…

I am exploring options just for fun. a used 3090 is around $900 on ebay. a used rtx 6000 ADA is around $5k 4 3090s are slower at inference and worse at training than 1 rtx 6000. 4x3090 would consume 1400W at load. Rtx 6000 would consume 300W at load. If you god forbid live in California and your power averages 45 cents per kwh, 4x3090 would be $1500+ more per year to operate than a single RTX 6000[0] [0] Back of the…

To make matters worse, the RTX3090 was released during the crypto craze and so a decent amount of the second hand market could contain overused GPUs that won’t last long, even if 3xxx to 4xxx performance difference is not that high, I would avoid the 3xxx series totally for resell value.

Re: 25L Portable NV-linked Dual 3090 LLM Rig

#58
post #29

OK, here's my quick critique of the article (having built a similar AM4-based system in 2023 for 2300€): 1) [I thought] The page is blocking cut & paste. Super annoying! 2) The exact mainboard is not specified exactly. There are 4 different boards called "ASUS ROG Strix X670E Gaming" and some of them only have one PCIe x16 slot. None of them can do PCIe x8 when using two GPUs. 3) The shopping link for the mainboard l…

> None of them can do PCIe x8 when using two GPUs.

Is that important for this workload? I thought most of the effort was spent processing data on the card rather than moving data on or off of it?

Re: 25L Portable NV-linked Dual 3090 LLM Rig

#59
post #50

Earlier quoted context omitted.

I have a similar setup as the author with 2x 3090s. The issue is not that it's slow. 20-30 tk/s is perfectly acceptable to me. The issue is that the quality of the models that I'm able to self-host pales in comparison to that of SOTA hosted models. They hallucinate more, don't follow prompts as well, and simply generate overall worse quality content. These are issues that plague all "AI" models, but they are particul…

On my 2x 3090s I am running glm4.5 air q1 and it runs at ~300pp and 20/30 tk/s works pretty well with roo code on vscode, rarely misses tool calls and produces decent quality code. I also tried to use it with claude code with claude code router and it's pretty fast. Roo code uses bigger contexts, so it's quite slower than claude code in general, but I like the workflow better. this is my snippet for llama-swap ``` mo…

Thanks, but I find it hard to believe that a Q1 model would produce decent results.

I see that the Q2 version is around 42GB, which might be doable on 2x 3090s, even if some of it spills over to CPU/RAM. Have you tried Q2?

Re: 25L Portable NV-linked Dual 3090 LLM Rig

#60
post #50

I built a similar system, meanwhile I've sold one of the RTX 3090's. Local inference is fun and feels liberating, but it's also slow, and once I was used to the immense power of the giant hosted models, the fun quickly disappeared. I've kept a single GPU to still be able to play a bit with light local models, but not anymore for serious use.

I have a similar setup as the author with 2x 3090s. The issue is not that it's slow. 20-30 tk/s is perfectly acceptable to me. The issue is that the quality of the models that I'm able to self-host pales in comparison to that of SOTA hosted models. They hallucinate more, don't follow prompts as well, and simply generate overall worse quality content. These are issues that plague all "AI" models, but they are particul…

> behemoth 100B+ parameter models, but to run those I would need to invest much more into this hobby than I'm willing to do.

Have you tried newer MoE models with llama.cpp's recent '--n-cpu-moe' option to offload MoE layers to the CPU? I can run gpt-oss-120b (5.1B active) on my 4080 and get a usable ~20 tk/s. Had to upgrade my system RAM, but that's easier. https://github.com/ggml-org/llama.cpp/discussions/15396 has a bit on getting that running

Post reply on HN