Earlier quoted context omitted.
[flagged]
A single 3090 will deliver more tflops than the m2 ultra.
Serving AI from the Basement – 192GB of VRAM Setup
11–20 of 279 posts
Re: Serving AI from the Basement – 192GB of VRAM Setup
#12Are you intending to use the capacity all for yourself or rent it out to others?
Re: Serving AI from the Basement – 192GB of VRAM Setup
#13Re: Serving AI from the Basement – 192GB of VRAM Setup
#14Re: Serving AI from the Basement – 192GB of VRAM Setup
#15Re: Serving AI from the Basement – 192GB of VRAM Setup
#16I hope this guy posts updates.
Re: Serving AI from the Basement – 192GB of VRAM Setup
#17this is why we need an actual AI blockchain, so we can donate GPU and earn rewards for the p2p api calls using the distributed model.
Is a blockchain needed to sell unused GPU capacity?
Re: Serving AI from the Basement – 192GB of VRAM Setup
#18Re: Serving AI from the Basement – 192GB of VRAM Setup
#19Earlier quoted context omitted.
[flagged]
A single 3090 will deliver more tflops than the m2 ultra.
Honestly though I'd be curious to see a cost analysis of Apple vs. Nvidia for commercial batched inference. The Nvidia system can obviously spit out more tokens/s but for the same price you could have multiple Mac Studios running the same model (and users would be dispatched to one of them).
Re: Serving AI from the Basement – 192GB of VRAM Setup
#20You could just buy a Mac Studio for 6500 USD, have 192 GB of unified RAM and have way less power consumption.
Are people running llama 3.1 405B on them?
I'm currently using reflection:70b_q4 which does a very good job in my opinion. It generates with 5.5 tokens/s for the response, which is just about my reading speed.
edit: I usually dont run larger models (q6) because of the speed. I'd guess a 405B model would just be awfully slow.