Earlier quoted context omitted.
Are people running llama 3.1 405B on them?
I'm running 70B models (usually in q4 .. q5_k_m, but possible to q6) on my 96Gbyte Macbook Pro with M2-Max (12 cpu cores, 38 gpu cores). This also leaves me with plenty of ram for other purposes. I'm currently using reflection:70b_q4 which does a very good job in my opinion. It generates with 5.5 tokens/s for the response, which is just about my reading speed. edit: I usually dont run larger models (q6) because of th…
Serving AI from the Basement – 192GB of VRAM Setup
71–80 of 279 posts
Re: Serving AI from the Basement – 192GB of VRAM Setup
#72Earlier quoted context omitted.
The main thing stopping me from going beyond 2x 4090’s in my home lab is power. Anything around ~2k watts on a single circuit breaker is likely to flip it, and that’s before you get to the costs involved of drawing that much power for multiple days of a training run. How did you navigate that in a (presumably) residential setting?
Not OP, but my current home had a dedicated 50A/240V circuit because the previous owner did glass work and had a massive electric kiln. I can't imagine it was cheap to install, but I've used it for beefy, energy hungry servers in the past. Which is all to say its possible in a residential setting, just probably expensive.
Re: Serving AI from the Basement – 192GB of VRAM Setup
#73Earlier quoted context omitted.
The main thing stopping me from going beyond 2x 4090’s in my home lab is power. Anything around ~2k watts on a single circuit breaker is likely to flip it, and that’s before you get to the costs involved of drawing that much power for multiple days of a training run. How did you navigate that in a (presumably) residential setting?
I can't believe a group of engineers are so afraid of residential power. It is not expensive, nor is it highly technical. It's not like we're factoring in latency and crosstalk... Read a quick howto, cruise into Home Depot and grab some legos off the shelf. Far easier to figure out than executing "hello world" without domain expertise.
Re: Serving AI from the Basement – 192GB of VRAM Setup
#74I thought I was balling with my dual 3090 with nvlink. I haven’t quite yet figured out what to do with 48GB VRAM yet. I hope this guy posts updates.
Re: Serving AI from the Basement – 192GB of VRAM Setup
#75> And who knows, maybe someone will look back on my work and be like “haha, remember when we thought 192GB of VRAM was a lot?” I wonder if this will happen. It's already really hard to buy big HDDs for my NAS because nobody buys external drives anymore. So the pricing has gone up a lot for the prosumer. I expect something similar to happen to AI. The big cloud parties are all big leaders on LLMs and their goal is to…
IME 20tb drives are easy to find.
I don't think the clouds have access to bigger drives or anything.
Similarly, we can buy 8x A100s, they're just fundamentally expensive whether you're a business or not.
There doesn't seem to be any "wall" up like there used to be with proprietary hardware.
Re: Serving AI from the Basement – 192GB of VRAM Setup
#76Earlier quoted context omitted.
The M2 Ultra doesn't require doing electrical work on your house like this 8x 3090 setup did though.
1x 3090 would be faster than the mac, cheaper than the mac, and run on a normal residential circuit. How many breakers would you need for a cluster of 20 macs running at full throttle to match the perf of OP's build?
This is exactly like when the AMD fanboys got a burr up their ass about the “$50k Mac Pro” with 2tb of memory… when you could the same thing with a threadripper with 256gb of memory for $5k, and it’s just as fast in Cinebench!
https://old.reddit.com/r/Amd/comments/f1a0qp/15000_mac_pro_v...
your gaming scores on your 3090 with 24gb are just as irrelevant to this 200gb workload as the threadripper is to those Mac Pro workloads lol
Re: Serving AI from the Basement – 192GB of VRAM Setup
#77Hey guys, this is something I have been intending to share here for a while. This setup took me some time to plan and put together, and then some more time to explore the software part of things and the possibilities that came with it. Part of the main reason I built this was data privacy, I do not want to hand over my private data to any company to further train their closed weight models; and given the recent drop…
Re: Serving AI from the Basement – 192GB of VRAM Setup
#78Earlier quoted context omitted.
I can't believe a group of engineers are so afraid of residential power. It is not expensive, nor is it highly technical. It's not like we're factoring in latency and crosstalk... Read a quick howto, cruise into Home Depot and grab some legos off the shelf. Far easier to figure out than executing "hello world" without domain expertise.
People can and do die from misuses of electricity. Not a move-fast-and-break things kind of domain.
Re: Serving AI from the Basement – 192GB of VRAM Setup
#79> And who knows, maybe someone will look back on my work and be like “haha, remember when we thought 192GB of VRAM was a lot?” I wonder if this will happen. It's already really hard to buy big HDDs for my NAS because nobody buys external drives anymore. So the pricing has gone up a lot for the prosumer. I expect something similar to happen to AI. The big cloud parties are all big leaders on LLMs and their goal is to…
Re: Serving AI from the Basement – 192GB of VRAM Setup
#80> And who knows, maybe someone will look back on my work and be like “haha, remember when we thought 192GB of VRAM was a lot?” I wonder if this will happen. It's already really hard to buy big HDDs for my NAS because nobody buys external drives anymore. So the pricing has gone up a lot for the prosumer. I expect something similar to happen to AI. The big cloud parties are all big leaders on LLMs and their goal is to…
> It's already really hard to buy big HDDs for my NAS IME 20tb drives are easy to find. I don't think the clouds have access to bigger drives or anything. Similarly, we can buy 8x A100s, they're just fundamentally expensive whether you're a business or not. There doesn't seem to be any "wall" up like there used to be with proprietary hardware.
For me these prices are prohibitive. Just like the A100s are (though those are even more so of course).
The problem is the common consumer relying on the cloud so these kind of products become niches and lose volume. Also, the cloud providers don't pay what we do for a GPU or HDD. They buy them by the ten thousands and get deep discounts. That's why the RRPs which we do pay are highly inflated.