Live data from Hacker News

Serving AI from the Basement – 192GB of VRAM Setup

ahmadosman.com

51–60 of 279 posts

Re: Serving AI from the Basement – 192GB of VRAM Setup

#51

Earlier quoted context omitted.

You're interested in the different between a single CPU and 8 GPUs? A Ford fiesta vs a freight train.

One can be interested in the differences between a Ford Fiesta and a freight train…

Are you fucking kidding me? A single train car can weight 130 tons, a fiesta can carry maybe 500kg, its not even close. /s

Re: Serving AI from the Basement – 192GB of VRAM Setup

#52
post #4

How much do the NVLinks help in this case? Do you have a rough estimate of how much this cost? I'm curious since I just built my own 2x 3090 rig and I wondered about going EPYC for the potential to have more cards (stuck with AM5 for cheapness though). All in all I spent about $3500 for everything. I'm guessing this is closer to $12-15k? CPU is around $800 on eBay.

My reason for going Epyc was for Pcie lanes and cheaper enterprise SSDs via U.3/2. With AM5, you tap out the lanes with dual GPUs. Threadripper is preferable but Epyc is about 1/2 of the price or even better if you go last gen.

Re: Serving AI from the Basement – 192GB of VRAM Setup

#53

Earlier quoted context omitted.

You could for sure, but the nVidia setup described in this article would be many times faster at inference. So it’s a tradeoff between power consumption and performance. Also, modern GPUs are surprisingly good at throttling their power usage when not actively in use, just like CPUs. So while you need 3kW+ worth of PSU for an 8x3090 setup, it’s not going to be using anywhere near 3kW of power on average, unless you’re…

Can Reflection:70b work on them?

Pretty sure it'll work where any 70b model would, but it's probably not noticably better than Llama 3.1 70b if the reports I'm reading now are correct.[1]

[1]https://x.com/JJitsev/status/1832758733866222011

Re: Serving AI from the Basement – 192GB of VRAM Setup

#54
post #7

You could just buy a Mac Studio for 6500 USD, have 192 GB of unified RAM and have way less power consumption.

You could for sure, but the nVidia setup described in this article would be many times faster at inference. So it’s a tradeoff between power consumption and performance. Also, modern GPUs are surprisingly good at throttling their power usage when not actively in use, just like CPUs. So while you need 3kW+ worth of PSU for an 8x3090 setup, it’s not going to be using anywhere near 3kW of power on average, unless you’re…

Even if you are running it constantly, the per token power consumption is likely going to be in a similar range, not to mention you'd need 10+ macs for the throughput.

Re: Serving AI from the Basement – 192GB of VRAM Setup

#55

Earlier quoted context omitted.

[flagged]

He's got 8x3090s are you fucking kidding? Like is this some kind of AI reply? "Wow great post! I enjoy your valuable contributions. Can you tell me more about graphics cards and how they compare to other different types of computers? I am interested and eager to learn! :)"

While this reply might be a bit too harsh, I fully agree with the well warranted criticism of Apple fans chiming in on every AI Nvidia discussion with “but M chips have large amount of RAM and Apple says they’re amazing for AI”.

Re: Serving AI from the Basement – 192GB of VRAM Setup

#56

Hey guys, this is something I have been intending to share here for a while. This setup took me some time to plan and put together, and then some more time to explore the software part of things and the possibilities that came with it. Part of the main reason I built this was data privacy, I do not want to hand over my private data to any company to further train their closed weight models; and given the recent drop…

The main thing stopping me from going beyond 2x 4090’s in my home lab is power. Anything around ~2k watts on a single circuit breaker is likely to flip it, and that’s before you get to the costs involved of drawing that much power for multiple days of a training run. How did you navigate that in a (presumably) residential setting?

Not OP, but my current home had a dedicated 50A/240V circuit because the previous owner did glass work and had a massive electric kiln. I can't imagine it was cheap to install, but I've used it for beefy, energy hungry servers in the past.

Which is all to say its possible in a residential setting, just probably expensive.

Re: Serving AI from the Basement – 192GB of VRAM Setup

#57

Earlier quoted context omitted.

[flagged]

He's got 8x3090s are you fucking kidding? Like is this some kind of AI reply? "Wow great post! I enjoy your valuable contributions. Can you tell me more about graphics cards and how they compare to other different types of computers? I am interested and eager to learn! :)"

It's one thing to be an asshole, but you're also hilariously clueless.

Re: Serving AI from the Basement – 192GB of VRAM Setup

#58
post #9

Earlier quoted context omitted.

A single 3090 will deliver more tflops than the m2 ultra.

The M2 Ultra doesn't require doing electrical work on your house like this 8x 3090 setup did though.

Location specific issue. Higher voltage resolves this and a regular household socket will power this just fine. Then again, if you're spending that much money to get the job done, new wiring is probably one of the cheaper parts of the build.

Re: Serving AI from the Basement – 192GB of VRAM Setup

#59

this is why we need an actual AI blockchain, so we can donate GPU and earn rewards for the p2p api calls using the distributed model.

That's actually interesting. While crypto GPU mining is "purposeless" or arbitrary, would be way cooler if to GPU mine meant to chunk through computing tasks in a free/open queue (blockchain). Eventually there could be some tipping point where networks are fast enough and there are enough hosting participants it could be like a worldwide/free computing platform - not just for AI for anything.

This idea has been brought up tons of times by grifters aiming to pivot from Crypto to AI. The reason that GPUs are used for blockchains is to compute large numbers or proofs - which are truly useless but still verifiable so they can be distributed and rewarded. The free GPU compute idea misses this crucial point, so the blockchain part is (still) useless unless your aim is to waste GPU compute instead.

IRL all you need is a simple platform to pay and schedule jobs on other’s GPUs.

Post reply on HN