Live data from Hacker News

Serving AI from the Basement – 192GB of VRAM Setup

ahmadosman.com

21–30 of 279 posts

Re: Serving AI from the Basement – 192GB of VRAM Setup

#21

Hey guys, this is something I have been intending to share here for a while. This setup took me some time to plan and put together, and then some more time to explore the software part of things and the possibilities that came with it. Part of the main reason I built this was data privacy, I do not want to hand over my private data to any company to further train their closed weight models; and given the recent drop…

[flagged]

He's got 8x3090s are you fucking kidding? Like is this some kind of AI reply?

"Wow great post! I enjoy your valuable contributions. Can you tell me more about graphics cards and how they compare to other different types of computers? I am interested and eager to learn! :)"

Re: Serving AI from the Basement – 192GB of VRAM Setup

#22
post #9

Earlier quoted context omitted.

A single 3090 will deliver more tflops than the m2 ultra.

The M2 Ultra doesn't require doing electrical work on your house like this 8x 3090 setup did though.

Yeah cause it's doing like 50x less work.

Re: Serving AI from the Basement – 192GB of VRAM Setup

#23
post #9

Earlier quoted context omitted.

A single 3090 will deliver more tflops than the m2 ultra.

The M2 Ultra doesn't require doing electrical work on your house like this 8x 3090 setup did though.

You could power limit the 3090s to fit a standard 120V*20A = 2400W outlet if you really want to. The default power limit is 350W each so you'll only lose a little perf.

Also most rooms have multiple circuits. Just connect half the GPUs to each outlet. I already do this with my desktop PC because it has 2 PSUs.

Also most homes in the US have 30A*240V = 7200W dryer/stove outlets in the kitchen, laundry room, garage, etc. That's where you want to keep loud computers.

Re: Serving AI from the Basement – 192GB of VRAM Setup

#24

this is why we need an actual AI blockchain, so we can donate GPU and earn rewards for the p2p api calls using the distributed model.

That's actually interesting. While crypto GPU mining is "purposeless" or arbitrary, would be way cooler if to GPU mine meant to chunk through computing tasks in a free/open queue (blockchain).

Eventually there could be some tipping point where networks are fast enough and there are enough hosting participants it could be like a worldwide/free computing platform - not just for AI for anything.

Re: Serving AI from the Basement – 192GB of VRAM Setup

#26

Earlier quoted context omitted.

[flagged]

He's got 8x3090s are you fucking kidding? Like is this some kind of AI reply? "Wow great post! I enjoy your valuable contributions. Can you tell me more about graphics cards and how they compare to other different types of computers? I am interested and eager to learn! :)"

What? No I just don’t know the difference, sorry. I am interested in learning more about running 405b parameter models, which I believe you can do on a 192gb M series Mac.

The answer here is that the Nvidia system has much better performance. I’ve been focused on “can I even run the model” I didn’t think about the actual performance of the system.

Re: Serving AI from the Basement – 192GB of VRAM Setup

#27

What is the power draw under load/idle? Does it noticeably increase the room temperature? Given the surroundings (aka the huge pile of boxes behind the setup), curious if you could get away with just a couple of box fans instead of the array of case fans. Are you intending to use the capacity all for yourself or rent it out to others?

Box fans are surprisingly power hungry. You'd be better off using large 200mm PC fans. They're also a lot quieter

Re: Serving AI from the Basement – 192GB of VRAM Setup

#28
post #7

You could just buy a Mac Studio for 6500 USD, have 192 GB of unified RAM and have way less power consumption.

This is something people often say without even attempting to do a major AI task. If Mac Studio were that great they’d be sold out completely. It’s not even cost efficient for inference.

Re: Serving AI from the Basement – 192GB of VRAM Setup

#29
post #23

Earlier quoted context omitted.

The M2 Ultra doesn't require doing electrical work on your house like this 8x 3090 setup did though.

You could power limit the 3090s to fit a standard 120V*20A = 2400W outlet if you really want to. The default power limit is 350W each so you'll only lose a little perf. Also most rooms have multiple circuits. Just connect half the GPUs to each outlet. I already do this with my desktop PC because it has 2 PSUs. Also most homes in the US have 30A*240V = 7200W dryer/stove outlets in the kitchen, laundry room, garage, et…

Brb replacing my stove with a bunch of GPUs

Re: Serving AI from the Basement – 192GB of VRAM Setup

#30

Earlier quoted context omitted.

He's got 8x3090s are you fucking kidding? Like is this some kind of AI reply? "Wow great post! I enjoy your valuable contributions. Can you tell me more about graphics cards and how they compare to other different types of computers? I am interested and eager to learn! :)"

What? No I just don’t know the difference, sorry. I am interested in learning more about running 405b parameter models, which I believe you can do on a 192gb M series Mac. The answer here is that the Nvidia system has much better performance. I’ve been focused on “can I even run the model” I didn’t think about the actual performance of the system.

You're interested in the different between a single CPU and 8 GPUs? A Ford fiesta vs a freight train.
Post reply on HN