Live data from Hacker News

Serving AI from the Basement – 192GB of VRAM Setup

ahmadosman.com

241–250 of 279 posts

Re: Serving AI from the Basement – 192GB of VRAM Setup

#241
post #168

Earlier quoted context omitted.

> Typical crypto miner setup. Except not doing the sketchy x1 pcie lanes. That’s the part that makes nice LLM setups hard

Can you tell me what's sketchy about it? I have not had an issue with any one of the 12 extenders and bandwidth held well without any issues. Please explain if possible if LLM requires a different type of extender.

Eh perhaps poor choice of words.

It works fine for crypto but LLM performance is far more sensitive to bandwidth. You lose a ton of performance if you’ve got PCIe in the loop, never mind one lane pcie. That’s why nvlink is (was) a thing - trying to cut that out entirely.

Re: Serving AI from the Basement – 192GB of VRAM Setup

#242
post #199

Earlier quoted context omitted.

I just get 99C water from a tap next to my kitchen sink. Why do people still use kettles?

Because they don’t have a spare few grand for an instant hot plus installation

You’re off by an order of magnitude. They are a couple hundred bucks and an easy DIY job.

Re: Serving AI from the Basement – 192GB of VRAM Setup

#243

Earlier quoted context omitted.

You can run a setup of 8x 4090 GPUs using 4x 1200W 240V power supplies (preferably HP HSTNS-PD30 Platinum Series), with a collective use of just around 20-amps, meaning it can easily run on a single 240V 20-amp breaker. This should be easily doable in a home where you typically have a 100 to 200A main power panel. Running 4x 1200W power supplies 24 hours a day will consume 115.2 kWh per day. At an electricity rate of…

In Cali isn't now like 0.5 per kWh :P

Wow!

Re: Serving AI from the Basement – 192GB of VRAM Setup

#245
post #52
post #4

How much do the NVLinks help in this case? Do you have a rough estimate of how much this cost? I'm curious since I just built my own 2x 3090 rig and I wondered about going EPYC for the potential to have more cards (stuck with AM5 for cheapness though). All in all I spent about $3500 for everything. I'm guessing this is closer to $12-15k? CPU is around $800 on eBay.

My reason for going Epyc was for Pcie lanes and cheaper enterprise SSDs via U.3/2. With AM5, you tap out the lanes with dual GPUs. Threadripper is preferable but Epyc is about 1/2 of the price or even better if you go last gen.

I tried this w/ AM5, but realized that despite there theoretically being enough lanes for dual x16 PCI-e 4.0 GPUs, I couldn't find any motherboards that are actually configured this way, since dual-GPU is dead in consumer for gaming.

Re: Serving AI from the Basement – 192GB of VRAM Setup

#246

Earlier quoted context omitted.

Yes and they get deep discounts which we don't. Can be 40% or more! Of course the vendor can't make a profit with such discounts so they inflate the RRP. But we do end up paying that.

That’s the main problem is a market owned by enterprise customers. Consumers don’t matter, there is zero interest is competing for them, they’re too little. The discounts is a killer for example, well have to buy from a reseller each time, who of course will pocket a good proportion of the discount because there won’t be many resellers that sell to consumers… I have seen very large ent customers get 80% discount on h…

Not specific for GPUs but I believe some of those giant and deeply discounted buys are at/below typical cost because of volume. They allow the vendor to increase their OEM/manufacturing commits, or shift bins theyre long on, to improve the rest of their sales pipeline. Similar for very large last orders or all the remaining stock of a SKU which improves cash flow and turns over inventory. Its a very very different vendor relationship with things like defect rates, yield, and “warranty” turned in to price factors.

Re: Serving AI from the Basement – 192GB of VRAM Setup

#247
post #136

Earlier quoted context omitted.

PCIe bifurcation - so splitting one of the x16 slots into two x8 or similar.

Worth mentioning - this also cuts the available bandwidth to each card by 50%.

While you're technically correct, assuming you're using PCIe 4.0 or higher, the performance difference between x8 and x16 is practically zero.

Re: Serving AI from the Basement – 192GB of VRAM Setup

#249
post #7

You could just buy a Mac Studio for 6500 USD, have 192 GB of unified RAM and have way less power consumption.

I'm seeing this misunderstanding a lot recently. There's TWO components to putting together a viable machine learning rig: - Fitting models in memory - Inference / Training speed 8 x RTX 3090s will absolutely CRUSH a single Mac Studio in raw performance.

Crush by what factor?

Re: Serving AI from the Basement – 192GB of VRAM Setup

#250
post #241

Earlier quoted context omitted.

Can you tell me what's sketchy about it? I have not had an issue with any one of the 12 extenders and bandwidth held well without any issues. Please explain if possible if LLM requires a different type of extender.

Eh perhaps poor choice of words. It works fine for crypto but LLM performance is far more sensitive to bandwidth. You lose a ton of performance if you’ve got PCIe in the loop, never mind one lane pcie. That’s why nvlink is (was) a thing - trying to cut that out entirely.

Got it. I was planning to switch my miners to LLM farm. I will test and see how much of a difference it will make. Thanks.
Post reply on HN