Live data from Hacker News

Was my $48K GPU server worth it?

rosmine.ai

471–480 of 480 posts

Re: Was my $48K GPU server worth it?

#472

Earlier quoted context omitted.

M5 pro 48GB should be good and future proof

Last time I looked into it, I realized how severely it's limited by memory bandwidth. Only the M5 Ultra compares to a dedicated graphics card, but it still falls short.

M5 Pro is not that expensive and allows for model below 35GB to be used which is a lot of models. And has a boost with neural engines, its not just the memory bandwidth that makes the speed

Re: Was my $48K GPU server worth it?

#473

Earlier quoted context omitted.

M5 pro 48GB should be good and future proof

If you buy Mac get at least 256GB ram otherwise just buy a bunch of nvidia cards. It really does not make sense otherwise if you are looking for performance / $. The mac (studio) is unique as it has more ram than the alternatives(I.e consumer nvidia cards or spark stuff) so it can fit bigger models but otherwise its performance is worse.

I’d rather use a 128GB mac which is super flexible than always configuring a desktop with a huge GPU.

Unless you’re doing others things with that gpu such as gaming, video, video AI etc.

Re: Was my $48K GPU server worth it?

#474
post #464

Earlier quoted context omitted.

Spot pricing and instance availability don’t apply to on metal hosting. You’d have your own machine dedicated to your own use only, at a locked in price.

Wait, I went and looked far and wide today and I can’t find anything with ~100GB of VRAM that isn’t $20k a year, what am I missing

That's too small. I was quoted a machine with ~1.5TB of vram, for $10k/mo. This was the minimum node size in the AI data center I was talking with -- they don't make smaller nodes that you can lease as bare metal. There is no public pricing though, and you had to know people to get in.

Re: Was my $48K GPU server worth it?

#475

Earlier quoted context omitted.

Does anyone here have experience running large models in a multi-GPU setup with several RTX 6000s in a high-concurrency regime and with large context lengths? (something like Deepseek 4 Flash, Minimax 2.7 etc.) For what it's worth, I've been seeing ~100 tps with 4-bit MiniMax 2.7 on two RTX 6000 boards, just running under llama-server without any optimization effort at all. I have no serious long-context experience w…

With that amount of memory can you run 4-bit DeepSeek 4 Flash? It is way more efficient in the KV cache department so may be worth a try

I haven't looked into DS4 yet but based on antirez's results on 128 GB Macbooks, it shouldn't be a problem to run it on a pair of RTX6000 Pros.

Also see https://www.reddit.com/r/LocalLLaMA/comments/1sv649s/to_run_... .

Re: Was my $48K GPU server worth it?

#476

Earlier quoted context omitted.

Because buying Macs is not about performance, its about feeling like you are rich. That money could have been spent on way more bang/buck performance in the form of a set of 4 graphics cards. Also I would probably put the odds 70:30 that Apple marketing is astroturfing on HN from the amount of posts about running llms on Macbooks, because in reality, the inference speed of any decent llm is unusable on a Macbook desp…

40-80 tok/s is unusable to you? Ok. If you like having a box with 8-12 fans blasting hot air and noise into your office all day, nobody's stopping you.

If you are not being paid by Apple, I feel sorry for you. Cause that means you are so bought into the cult that you are delusional.

the 40-80 tok/sec is only for initial prompt processing, and with the "medium" models, like Qwen3.6:27b. The actual token generation is in the 10 token/second Thats very slow. And your Macbook pro will stop being a LAP-top, because it will get very warm.

Meanwhile, my 2x3090s happily crank out ~100 tok/sec generation. Oh and I can run 100 tok/sec on my phone as well, because I can just access ollama on my home desktop over ssh from termux.

Re: Was my $48K GPU server worth it?

#479

Earlier quoted context omitted.

40-80 tok/s is unusable to you? Ok. If you like having a box with 8-12 fans blasting hot air and noise into your office all day, nobody's stopping you.

If you are not being paid by Apple, I feel sorry for you. Cause that means you are so bought into the cult that you are delusional. the 40-80 tok/sec is only for initial prompt processing, and with the "medium" models, like Qwen3.6:27b. The actual token generation is in the 10 token/second Thats very slow. And your Macbook pro will stop being a LAP-top, because it will get very warm. Meanwhile, my 2x3090s happily cra…

Please, allow me to show you how far I can pee. It's several feet more than you can and a much mightier stream to boot.

Re: Was my $48K GPU server worth it?

#480
post #236

Earlier quoted context omitted.

Because buying Macs is not about performance, its about feeling like you are rich. That money could have been spent on way more bang/buck performance in the form of a set of 4 graphics cards. Also I would probably put the odds 70:30 that Apple marketing is astroturfing on HN from the amount of posts about running llms on Macbooks, because in reality, the inference speed of any decent llm is unusable on a Macbook desp…

Or it could have had way more bang/buck by feeding a family of real brains for a year or two

[deleted]
Post reply on HN