Was my $48K GPU server worth it?
471–480 of 480 posts
Re: Was my $48K GPU server worth it?
#472Earlier quoted context omitted.
M5 pro 48GB should be good and future proof
Last time I looked into it, I realized how severely it's limited by memory bandwidth. Only the M5 Ultra compares to a dedicated graphics card, but it still falls short.
Re: Was my $48K GPU server worth it?
#473Earlier quoted context omitted.
M5 pro 48GB should be good and future proof
If you buy Mac get at least 256GB ram otherwise just buy a bunch of nvidia cards. It really does not make sense otherwise if you are looking for performance / $. The mac (studio) is unique as it has more ram than the alternatives(I.e consumer nvidia cards or spark stuff) so it can fit bigger models but otherwise its performance is worse.
Unless you’re doing others things with that gpu such as gaming, video, video AI etc.
Re: Was my $48K GPU server worth it?
#474Earlier quoted context omitted.
Spot pricing and instance availability don’t apply to on metal hosting. You’d have your own machine dedicated to your own use only, at a locked in price.
Wait, I went and looked far and wide today and I can’t find anything with ~100GB of VRAM that isn’t $20k a year, what am I missing
Re: Was my $48K GPU server worth it?
#475Earlier quoted context omitted.
Does anyone here have experience running large models in a multi-GPU setup with several RTX 6000s in a high-concurrency regime and with large context lengths? (something like Deepseek 4 Flash, Minimax 2.7 etc.) For what it's worth, I've been seeing ~100 tps with 4-bit MiniMax 2.7 on two RTX 6000 boards, just running under llama-server without any optimization effort at all. I have no serious long-context experience w…
With that amount of memory can you run 4-bit DeepSeek 4 Flash? It is way more efficient in the KV cache department so may be worth a try
Also see https://www.reddit.com/r/LocalLLaMA/comments/1sv649s/to_run_... .
Re: Was my $48K GPU server worth it?
#476Earlier quoted context omitted.
Because buying Macs is not about performance, its about feeling like you are rich. That money could have been spent on way more bang/buck performance in the form of a set of 4 graphics cards. Also I would probably put the odds 70:30 that Apple marketing is astroturfing on HN from the amount of posts about running llms on Macbooks, because in reality, the inference speed of any decent llm is unusable on a Macbook desp…
40-80 tok/s is unusable to you? Ok. If you like having a box with 8-12 fans blasting hot air and noise into your office all day, nobody's stopping you.
the 40-80 tok/sec is only for initial prompt processing, and with the "medium" models, like Qwen3.6:27b. The actual token generation is in the 10 token/second Thats very slow. And your Macbook pro will stop being a LAP-top, because it will get very warm.
Meanwhile, my 2x3090s happily crank out ~100 tok/sec generation. Oh and I can run 100 tok/sec on my phone as well, because I can just access ollama on my home desktop over ssh from termux.
Re: Was my $48K GPU server worth it?
#477Re: Was my $48K GPU server worth it?
#478Re: Was my $48K GPU server worth it?
#479Earlier quoted context omitted.
40-80 tok/s is unusable to you? Ok. If you like having a box with 8-12 fans blasting hot air and noise into your office all day, nobody's stopping you.
If you are not being paid by Apple, I feel sorry for you. Cause that means you are so bought into the cult that you are delusional. the 40-80 tok/sec is only for initial prompt processing, and with the "medium" models, like Qwen3.6:27b. The actual token generation is in the 10 token/second Thats very slow. And your Macbook pro will stop being a LAP-top, because it will get very warm. Meanwhile, my 2x3090s happily cra…
Re: Was my $48K GPU server worth it?
#480Earlier quoted context omitted.
Because buying Macs is not about performance, its about feeling like you are rich. That money could have been spent on way more bang/buck performance in the form of a set of 4 graphics cards. Also I would probably put the odds 70:30 that Apple marketing is astroturfing on HN from the amount of posts about running llms on Macbooks, because in reality, the inference speed of any decent llm is unusable on a Macbook desp…
Or it could have had way more bang/buck by feeding a family of real brains for a year or two