Live data from Hacker News

Serving AI from the Basement – 192GB of VRAM Setup

ahmadosman.com

81–90 of 279 posts

Re: Serving AI from the Basement – 192GB of VRAM Setup

#81

Earlier quoted context omitted.

He's got 8x3090s are you fucking kidding? Like is this some kind of AI reply? "Wow great post! I enjoy your valuable contributions. Can you tell me more about graphics cards and how they compare to other different types of computers? I am interested and eager to learn! :)"

It's one thing to be an asshole, but you're also hilariously clueless.

Yeah, because an M2 is in the same ballpark as 8 GPUs. Yes, you can use CPU now but it's not even close to this setup. This is hackernews. I know we're supposed to be nice and this isn't reddit, but comments like parent are ridiculous and for sure don't add to the discussion any more than mine do.

Re: Serving AI from the Basement – 192GB of VRAM Setup

#82
post #23

Earlier quoted context omitted.

You could power limit the 3090s to fit a standard 120V*20A = 2400W outlet if you really want to. The default power limit is 350W each so you'll only lose a little perf. Also most rooms have multiple circuits. Just connect half the GPUs to each outlet. I already do this with my desktop PC because it has 2 PSUs. Also most homes in the US have 30A*240V = 7200W dryer/stove outlets in the kitchen, laundry room, garage, et…

Brb replacing my stove with a bunch of GPUs

You can cook on the GPUs from now on.

Re: Serving AI from the Basement – 192GB of VRAM Setup

#83
post #59

Earlier quoted context omitted.

This idea has been brought up tons of times by grifters aiming to pivot from Crypto to AI. The reason that GPUs are used for blockchains is to compute large numbers or proofs - which are truly useless but still verifiable so they can be distributed and rewarded. The free GPU compute idea misses this crucial point, so the blockchain part is (still) useless unless your aim is to waste GPU compute instead. IRL all you n…

folding@home predates Bitcoin by eight years. the concept isn't inherent to grifters

Folding at home does not use a blockchain, further proving non-grifters don’t need it. That was the point being discussed, not distributed computing as a concept.

Re: Serving AI from the Basement – 192GB of VRAM Setup

#84

Earlier quoted context omitted.

One can be interested in the differences between a Ford Fiesta and a freight train…

Are you fucking kidding me? A single train car can weight 130 tons, a fiesta can carry maybe 500kg, its not even close. /s

How is that even sarcasm

Re: Serving AI from the Basement – 192GB of VRAM Setup

#85

> And who knows, maybe someone will look back on my work and be like “haha, remember when we thought 192GB of VRAM was a lot?” I wonder if this will happen. It's already really hard to buy big HDDs for my NAS because nobody buys external drives anymore. So the pricing has gone up a lot for the prosumer. I expect something similar to happen to AI. The big cloud parties are all big leaders on LLMs and their goal is to…

It isn't that cloud providers want to shut us out, it is that nVidia wants to relegate AI capable cards to the high end enterprise tier. So far in 2024 they have made $10.44b in revenue from the gaming market, and over $47.5b in the datacenter market, and I would bet that there is much less profit in gaming. In order to keep the market segmented they stopped putting nvlink on gaming cards and have capped VRAM at 24GB for the highest end GPUs (3090 and 4090) and it doesn't look much better for the upcoming 5090. I don't blame them, they are a profit-maximizing corporation after all, but if anything is to be done about making large AI models practical for hobbyists, start with nVidia.

That said, I really don't think that the way forward for hobbyists is maxing VRAM. Small models are becoming much more capable and accelerators are a possibility, and there may not be a need for a person to run a 70billion parameter model in memory at all when there are MoEs like Mixtral and small capable models like phi.

Re: Serving AI from the Basement – 192GB of VRAM Setup

#86
post #23

Earlier quoted context omitted.

The M2 Ultra doesn't require doing electrical work on your house like this 8x 3090 setup did though.

You could power limit the 3090s to fit a standard 120V*20A = 2400W outlet if you really want to. The default power limit is 350W each so you'll only lose a little perf. Also most rooms have multiple circuits. Just connect half the GPUs to each outlet. I already do this with my desktop PC because it has 2 PSUs. Also most homes in the US have 30A*240V = 7200W dryer/stove outlets in the kitchen, laundry room, garage, et…

A standard outlet is 15A, and you're only allowed to use 80% continuously. So now you're dropping to 150W each.

Using two circuits is viable.

Re: Serving AI from the Basement – 192GB of VRAM Setup

#87

> And who knows, maybe someone will look back on my work and be like “haha, remember when we thought 192GB of VRAM was a lot?” I wonder if this will happen. It's already really hard to buy big HDDs for my NAS because nobody buys external drives anymore. So the pricing has gone up a lot for the prosumer. I expect something similar to happen to AI. The big cloud parties are all big leaders on LLMs and their goal is to…

The cloud companies do not make the hardware, they buy it like the rest of us. They are just going to be almost the entirety of the market, so naturally the products will built and priced with that market in mind.

Yes and they get deep discounts which we don't. Can be 40% or more!

Of course the vendor can't make a profit with such discounts so they inflate the RRP. But we do end up paying that.

Re: Serving AI from the Basement – 192GB of VRAM Setup

#89
post #52
post #4

How much do the NVLinks help in this case? Do you have a rough estimate of how much this cost? I'm curious since I just built my own 2x 3090 rig and I wondered about going EPYC for the potential to have more cards (stuck with AM5 for cheapness though). All in all I spent about $3500 for everything. I'm guessing this is closer to $12-15k? CPU is around $800 on eBay.

My reason for going Epyc was for Pcie lanes and cheaper enterprise SSDs via U.3/2. With AM5, you tap out the lanes with dual GPUs. Threadripper is preferable but Epyc is about 1/2 of the price or even better if you go last gen.

Why do you need such high cross card bandwidth for inference? Are you hosting for a lot of users at once?

Re: Serving AI from the Basement – 192GB of VRAM Setup

#90

Hey guys, this is something I have been intending to share here for a while. This setup took me some time to plan and put together, and then some more time to explore the software part of things and the possibilities that came with it. Part of the main reason I built this was data privacy, I do not want to hand over my private data to any company to further train their closed weight models; and given the recent drop…

The main thing stopping me from going beyond 2x 4090’s in my home lab is power. Anything around ~2k watts on a single circuit breaker is likely to flip it, and that’s before you get to the costs involved of drawing that much power for multiple days of a training run. How did you navigate that in a (presumably) residential setting?

[deleted]
Post reply on HN