Live data from Hacker News

Serving AI from the Basement – 192GB of VRAM Setup

ahmadosman.com

31–40 of 279 posts

Re: Serving AI from the Basement – 192GB of VRAM Setup

#32
post #9

Earlier quoted context omitted.

A single 3090 will deliver more tflops than the m2 ultra.

Yep. The 8x3090 setup should be at least 10 times faster than the Mac.

Yes though for inference memory bandwidth is more important than tflops. Not sure how Apple compares in that regard. The OP does train models too though which is more compute heavy.

Re: Serving AI from the Basement – 192GB of VRAM Setup

#33

Earlier quoted context omitted.

What? No I just don’t know the difference, sorry. I am interested in learning more about running 405b parameter models, which I believe you can do on a 192gb M series Mac. The answer here is that the Nvidia system has much better performance. I’ve been focused on “can I even run the model” I didn’t think about the actual performance of the system.

You're interested in the different between a single CPU and 8 GPUs? A Ford fiesta vs a freight train.

One can be interested in the differences between a Ford Fiesta and a freight train…

Re: Serving AI from the Basement – 192GB of VRAM Setup

#35
An adjacent project for 8 GPUs could convert used 4K monitors into a borderless mini-wall of pixels, for local video composition with rendered and/or AI-generated backgrounds, https://theasc.com/articles/the-mandalorian

> the heir to rear projection — a dynamic, real-time, photo-real background played back on a massive LED video wall and ceiling, which not only provided the pixel-accurate representation of exotic background content, but was also rendered with correct camera positional data.. “We take objects that the art department have created and we employ photogrammetry on each item to get them into the game engine”

Re: Serving AI from the Basement – 192GB of VRAM Setup

#36
post #23

Earlier quoted context omitted.

You could power limit the 3090s to fit a standard 120V*20A = 2400W outlet if you really want to. The default power limit is 350W each so you'll only lose a little perf. Also most rooms have multiple circuits. Just connect half the GPUs to each outlet. I already do this with my desktop PC because it has 2 PSUs. Also most homes in the US have 30A*240V = 7200W dryer/stove outlets in the kitchen, laundry room, garage, et…

Brb replacing my stove with a bunch of GPUs

Inference Oven

Re: Serving AI from the Basement – 192GB of VRAM Setup

#37
I have a similar one with 4090s. Very cool. Yours is nicer than mine where I've let the 4090s rattle around a bit.

I haven't had enough time to find a way to split inference which is what I'm most interested in. Yours is also much better with the 1600 W supply. I have a hodge podge.

Re: Serving AI from the Basement – 192GB of VRAM Setup

#38

Earlier quoted context omitted.

What? No I just don’t know the difference, sorry. I am interested in learning more about running 405b parameter models, which I believe you can do on a 192gb M series Mac. The answer here is that the Nvidia system has much better performance. I’ve been focused on “can I even run the model” I didn’t think about the actual performance of the system.

You're interested in the different between a single CPU and 8 GPUs? A Ford fiesta vs a freight train.

A single SoC, which includes a GPU (two GPUs, kinda).

Re: Serving AI from the Basement – 192GB of VRAM Setup

#39

Hey guys, this is something I have been intending to share here for a while. This setup took me some time to plan and put together, and then some more time to explore the software part of things and the possibilities that came with it. Part of the main reason I built this was data privacy, I do not want to hand over my private data to any company to further train their closed weight models; and given the recent drop…

Do you run this 24/7? What is your cost of electricity per kilowatt hour and what is the cost of this setup per month?

I have a much smaller setup than the author - a quarter the GPUs and RAM - and I was surprised to find it draws 300W at idle

Re: Serving AI from the Basement – 192GB of VRAM Setup

#40
post #9

Earlier quoted context omitted.

A single 3090 will deliver more tflops than the m2 ultra.

The M2 Ultra doesn't require doing electrical work on your house like this 8x 3090 setup did though.

Sure, neither does just using ChatGPT, but we don't do things because they are easy, we do them because they are difficult.
Post reply on HN