Live data from Hacker News

Serving AI from the Basement – 192GB of VRAM Setup

ahmadosman.com

71–80 of 279 posts

Re: Serving AI from the Basement – 192GB of VRAM Setup

#71
post #8

Earlier quoted context omitted.

Are people running llama 3.1 405B on them?

I'm running 70B models (usually in q4 .. q5_k_m, but possible to q6) on my 96Gbyte Macbook Pro with M2-Max (12 cpu cores, 38 gpu cores). This also leaves me with plenty of ram for other purposes. I'm currently using reflection:70b_q4 which does a very good job in my opinion. It generates with 5.5 tokens/s for the response, which is just about my reading speed. edit: I usually dont run larger models (q6) because of th…

Not going to work for training from scratch which is what the author is doing.

Re: Serving AI from the Basement – 192GB of VRAM Setup

#72
post #56

Earlier quoted context omitted.

The main thing stopping me from going beyond 2x 4090’s in my home lab is power. Anything around ~2k watts on a single circuit breaker is likely to flip it, and that’s before you get to the costs involved of drawing that much power for multiple days of a training run. How did you navigate that in a (presumably) residential setting?

Not OP, but my current home had a dedicated 50A/240V circuit because the previous owner did glass work and had a massive electric kiln. I can't imagine it was cheap to install, but I've used it for beefy, energy hungry servers in the past. Which is all to say its possible in a residential setting, just probably expensive.

Yes, or something like a residential aircon heatpump will need a 40a circuit too. Car charging usually has a 30a. Electric oven is usually 40a. There’s lots of stuff that uses that sort of power residentially

Re: Serving AI from the Basement – 192GB of VRAM Setup

#73
post #70

Earlier quoted context omitted.

The main thing stopping me from going beyond 2x 4090’s in my home lab is power. Anything around ~2k watts on a single circuit breaker is likely to flip it, and that’s before you get to the costs involved of drawing that much power for multiple days of a training run. How did you navigate that in a (presumably) residential setting?

I can't believe a group of engineers are so afraid of residential power. It is not expensive, nor is it highly technical. It's not like we're factoring in latency and crosstalk... Read a quick howto, cruise into Home Depot and grab some legos off the shelf. Far easier to figure out than executing "hello world" without domain expertise.

People can and do die from misuses of electricity. Not a move-fast-and-break things kind of domain.

Re: Serving AI from the Basement – 192GB of VRAM Setup

#75

> And who knows, maybe someone will look back on my work and be like “haha, remember when we thought 192GB of VRAM was a lot?” I wonder if this will happen. It's already really hard to buy big HDDs for my NAS because nobody buys external drives anymore. So the pricing has gone up a lot for the prosumer. I expect something similar to happen to AI. The big cloud parties are all big leaders on LLMs and their goal is to…

> It's already really hard to buy big HDDs for my NAS

IME 20tb drives are easy to find.

I don't think the clouds have access to bigger drives or anything.

Similarly, we can buy 8x A100s, they're just fundamentally expensive whether you're a business or not.

There doesn't seem to be any "wall" up like there used to be with proprietary hardware.

Re: Serving AI from the Basement – 192GB of VRAM Setup

#76

Earlier quoted context omitted.

The M2 Ultra doesn't require doing electrical work on your house like this 8x 3090 setup did though.

1x 3090 would be faster than the mac, cheaper than the mac, and run on a normal residential circuit. How many breakers would you need for a cluster of 20 macs running at full throttle to match the perf of OP's build?

A single 3090 won’t even fit the model. OP is talking about running a 405B model, it needs close to 200gb of memory just to open it which is why we’re talking about mac studio and it’s 192gb unified memory.

This is exactly like when the AMD fanboys got a burr up their ass about the “$50k Mac Pro” with 2tb of memory… when you could the same thing with a threadripper with 256gb of memory for $5k, and it’s just as fast in Cinebench!

https://old.reddit.com/r/Amd/comments/f1a0qp/15000_mac_pro_v...

your gaming scores on your 3090 with 24gb are just as irrelevant to this 200gb workload as the threadripper is to those Mac Pro workloads lol

Re: Serving AI from the Basement – 192GB of VRAM Setup

#77

Hey guys, this is something I have been intending to share here for a while. This setup took me some time to plan and put together, and then some more time to explore the software part of things and the possibilities that came with it. Part of the main reason I built this was data privacy, I do not want to hand over my private data to any company to further train their closed weight models; and given the recent drop…

Amazing setup. I have the capability to design, fabricate, and powder coat sheet metal. I would love to collaborate on designing and fabricating a cool enclosure for this setup. Let me know if you're interested.

Re: Serving AI from the Basement – 192GB of VRAM Setup

#78
post #70

Earlier quoted context omitted.

I can't believe a group of engineers are so afraid of residential power. It is not expensive, nor is it highly technical. It's not like we're factoring in latency and crosstalk... Read a quick howto, cruise into Home Depot and grab some legos off the shelf. Far easier to figure out than executing "hello world" without domain expertise.

People can and do die from misuses of electricity. Not a move-fast-and-break things kind of domain.

You only "break" once...

Re: Serving AI from the Basement – 192GB of VRAM Setup

#79

> And who knows, maybe someone will look back on my work and be like “haha, remember when we thought 192GB of VRAM was a lot?” I wonder if this will happen. It's already really hard to buy big HDDs for my NAS because nobody buys external drives anymore. So the pricing has gone up a lot for the prosumer. I expect something similar to happen to AI. The big cloud parties are all big leaders on LLMs and their goal is to…

The cloud companies do not make the hardware, they buy it like the rest of us. They are just going to be almost the entirety of the market, so naturally the products will built and priced with that market in mind.

Re: Serving AI from the Basement – 192GB of VRAM Setup

#80

> And who knows, maybe someone will look back on my work and be like “haha, remember when we thought 192GB of VRAM was a lot?” I wonder if this will happen. It's already really hard to buy big HDDs for my NAS because nobody buys external drives anymore. So the pricing has gone up a lot for the prosumer. I expect something similar to happen to AI. The big cloud parties are all big leaders on LLMs and their goal is to…

> It's already really hard to buy big HDDs for my NAS IME 20tb drives are easy to find. I don't think the clouds have access to bigger drives or anything. Similarly, we can buy 8x A100s, they're just fundamentally expensive whether you're a business or not. There doesn't seem to be any "wall" up like there used to be with proprietary hardware.

They are easy to find but extremely expensive. I used to pay below 200€ for a 14TB Seagate 8 years ago. That's now above 300. And the bigger ones are even more expensive.

For me these prices are prohibitive. Just like the A100s are (though those are even more so of course).

The problem is the common consumer relying on the cloud so these kind of products become niches and lose volume. Also, the cloud providers don't pay what we do for a GPU or HDD. They buy them by the ten thousands and get deep discounts. That's why the RRPs which we do pay are highly inflated.

Post reply on HN