Live data from Hacker News

Serving AI from the Basement – 192GB of VRAM Setup

ahmadosman.com

171–180 of 279 posts

Re: Serving AI from the Basement – 192GB of VRAM Setup

#171
post #167
post #124

Earlier quoted context omitted.

In the US, it’s fully legal to perform electric/plumbing/whatever work on your own home. If you screw it up and need to file a claim, insurance can’t deny the claim based solely on the fact that you performed the work yourself, even if you’re not a certified electrician/plumber/whatever. What you don't want to do is have an unlicensed friend work on your home, and vice versa. There are no legal protections, and the i…

Extraordinary claims require extraordinary evidence.

As with most regulations in the "US" I have a feeling the answer is really something like "Depending on the city and state you live in the answer lies somewhere between 'go nuts' and 'that could lead to criminal charges and you being liable for everything that happens to the house and your neighbors kitchen sink'".

Re: Serving AI from the Basement – 192GB of VRAM Setup

#172

Hey guys, this is something I have been intending to share here for a while. This setup took me some time to plan and put together, and then some more time to explore the software part of things and the possibilities that came with it. Part of the main reason I built this was data privacy, I do not want to hand over my private data to any company to further train their closed weight models; and given the recent drop…

The main thing stopping me from going beyond 2x 4090’s in my home lab is power. Anything around ~2k watts on a single circuit breaker is likely to flip it, and that’s before you get to the costs involved of drawing that much power for multiple days of a training run. How did you navigate that in a (presumably) residential setting?

I use a hair dryer that is a little bit more than 2kw, but I guess because of the 120V it would be a problem in the US.

16 amps x 120v = 1920W, it would probably trip after several minutes.

16 amps x 230v = 3680W, it wouldn't trip.

Re: Serving AI from the Basement – 192GB of VRAM Setup

#173
post #172

Earlier quoted context omitted.

The main thing stopping me from going beyond 2x 4090’s in my home lab is power. Anything around ~2k watts on a single circuit breaker is likely to flip it, and that’s before you get to the costs involved of drawing that much power for multiple days of a training run. How did you navigate that in a (presumably) residential setting?

I use a hair dryer that is a little bit more than 2kw, but I guess because of the 120V it would be a problem in the US. 16 amps x 120v = 1920W, it would probably trip after several minutes. 16 amps x 230v = 3680W, it wouldn't trip.

When my gf first came to Europe she brought her hairdryer from the US and plugged it in using an adapter that just reroutes the wires. She was unaware of the voltage difference (or thought the adapter would adjust it.) That thing started spewing fire pretty much immediately and I luckily quickly realized what she was doing and was able to pull the plug (I hadn't noticed that she brought her own hair dryer.) Luckily she wasn't pointing it at herself ...

Re: Serving AI from the Basement – 192GB of VRAM Setup

#174
post #124
post #105

Earlier quoted context omitted.

Add to that is that it is likely illegal to do yourself. Which of course has implications for insurance etc.

In the US, it’s fully legal to perform electric/plumbing/whatever work on your own home. If you screw it up and need to file a claim, insurance can’t deny the claim based solely on the fact that you performed the work yourself, even if you’re not a certified electrician/plumber/whatever. What you don't want to do is have an unlicensed friend work on your home, and vice versa. There are no legal protections, and the i…

In my jurisdiction I can certainly do the work but am under the same requirements to pull a permit and pass a provincial inspection. It very quickly becomes the most effective to have an electrician involved, maybe not for all the work but some of it. They're more that willing to review the work you do and talk about it. Think of it as pair coding - great opportunity to learn and they'll tell you when you've done a good job. (at least the ones I've found)

Re: Serving AI from the Basement – 192GB of VRAM Setup

#175

Earlier quoted context omitted.

What? No I just don’t know the difference, sorry. I am interested in learning more about running 405b parameter models, which I believe you can do on a 192gb M series Mac. The answer here is that the Nvidia system has much better performance. I’ve been focused on “can I even run the model” I didn’t think about the actual performance of the system.

It's kinda hard to believe that someone would stumble onto the landmine of AI performance comparison between Apple Silicon and Nvidia hardware. People are going to be rude because this kinda behavior is genuinely indistinguishable from bad-faith trolling. From benchmarks alone, you can easily tell that the performance-per-watt of any Mac Studio gets annihilated by a 4090: https://browser.geekbench.com/opencl-benchmar…

> It's kinda hard to believe that someone would stumble onto the landmine of AI performance comparison between Apple Silicon and Nvidia hardware.

I encourage you to update your beliefs about other people. I’m a very technical person, but I work in robotics closer to the hardware level - I design motor controllers and Linux motherboards and write firmware and platform level robotics stacks, but I’ve never done any work that required running inference in a professional capacity. I’ve played with machine learning, even collecting and hand labeling my own dataset and training a semantic segmentation network. But I’ve only ever had my little desktop with one Nvidia card to run it all. Back in the day, performance of CNNs was very important and I might have looked at benchmarks, but since the dawn of LLMs, my ability to run networks has been limited entirely by RAM constraints, not other factors like tokens per second. So when I heard that MacBooks have shared memory and can run large models with it, I started to notice that could be a (relatively) accessible way to run larger models. I can’t even remotely afford a $6k Mac any more than I could afford a $12k Nvidia cluster machine, so I never really got to the practical considerations of whether there would be any serious performance concerns. It has been idle thinking like “hmm I wonder how well that would work”.

So I asked the question. I said roughly “hey can someone explain why OP didn’t go with this cheaper solution”. The very simple answer is that it would be much slower and the performance per dollar would be 10x worse. Great! Question answered. All this rude incredulousness coming from people who cannot fathom that another person might not know the answer is really odd to me. I simply never even thought to check benchmarks because it was never a real consideration for me to buy a system.

Also the “#1 topic in the tech sector right now” funny in my circles people are talking about unions, AI compute exacerbating climate change, and AI being used to disenfranchise and make more precarious the tech working class. We all live in bubbles.

Re: Serving AI from the Basement – 192GB of VRAM Setup

#176

Earlier quoted context omitted.

What? No I just don’t know the difference, sorry. I am interested in learning more about running 405b parameter models, which I believe you can do on a 192gb M series Mac. The answer here is that the Nvidia system has much better performance. I’ve been focused on “can I even run the model” I didn’t think about the actual performance of the system.

You're interested in the different between a single CPU and 8 GPUs? A Ford fiesta vs a freight train.

Yeah. I can’t afford a freight train.

Re: Serving AI from the Basement – 192GB of VRAM Setup

#177
post #124
post #105

Earlier quoted context omitted.

Add to that is that it is likely illegal to do yourself. Which of course has implications for insurance etc.

In the US, it’s fully legal to perform electric/plumbing/whatever work on your own home. If you screw it up and need to file a claim, insurance can’t deny the claim based solely on the fact that you performed the work yourself, even if you’re not a certified electrician/plumber/whatever. What you don't want to do is have an unlicensed friend work on your home, and vice versa. There are no legal protections, and the i…

Depends on the state and municipality. Mine doesn't allow homeowners to pull electrical permits.

Re: Serving AI from the Basement – 192GB of VRAM Setup

#178

Earlier quoted context omitted.

It's one thing to be an asshole, but you're also hilariously clueless.

Yeah, because an M2 is in the same ballpark as 8 GPUs. Yes, you can use CPU now but it's not even close to this setup. This is hackernews. I know we're supposed to be nice and this isn't reddit, but comments like parent are ridiculous and for sure don't add to the discussion any more than mine do.

I simply didn’t know the answer and the response “it would be much slower” is a perfectly acceptable reply. I disagree that it was ridiculous to ask. I was curious and I wanted to know the answer and now I do. What is obvious to you is not obvious to other people, and you will get nowhere in life insulting people who are asking questions in good faith.

Re: Serving AI from the Basement – 192GB of VRAM Setup

#179

Earlier quoted context omitted.

I know it's a fraction of the size, but my 32GB studio gets wrecked by these types of tasks. My experience is that they're awesome computers in general, but not as good for AI as people expect. Running llama3.1 70B is brutal on this thing. Responses take minutes. Someone running the same model on 32GB of GPU memory seems to have far better results from what I've read.

You are probably swapping. On M3 max with similar memory bandwidth the output is around 4t/s which is normally on par with most people's reading speed. Try different quants.

I'm on an M2 max so I shouldn't be too far behind. I'm not actually sure how the model I'm using was quantized to be honest. I'll give it a try.

Re: Serving AI from the Basement – 192GB of VRAM Setup

#180

Earlier quoted context omitted.

You're interested in the different between a single CPU and 8 GPUs? A Ford fiesta vs a freight train.

Yeah. I can’t afford a freight train.

Keep an eye on the SV going out of business fire sales. Not all the AI kites will fly.

As for actual trains, they can be suprisingly affordable (to live in):

https://atrservices.com.au/product/sa-red-hen-416/ https://en.wikipedia.org/wiki/South_Australian_Railways_Redh...

and the freight rolling stock flatcars make great bridges (single or sectioned) with concrete pylons either end for farms - once the axles are shot they can go pretty damn cheap and the beds are good enough to roll a car or small truck over.

addendum: in case you miss fresh reply to old comment: https://news.ycombinator.com/item?id=41484529

Post reply on HN