Live data from Hacker News

Was my $48K GPU server worth it?

rosmine.ai

141–150 of 480 posts

Re: Was my $48K GPU server worth it?

#141
post #45
post #25

The idea is similar to maintaining on-prem vs cloud Cloud is optimized for development velocity but its nature of high margin business eventually makes on-prem more promising It could be too late but it might be worth looking into tax saving if you have a business. Depreciation of asset is a loss and may deduct your income. (I'm NOT a tax expert)

Cloud servers have cheaper electricity, the scale of industrial-level cooling, no issues for you (as a user) with hardware failure (ie you just use a different server; it's not your problem) and can amortize their cost by running 24x7. I've seen H100 computer hours for as little as $2. As the author notes, there are also electrical/wiring issues that cap how much compute gear you can run in a space not designed for i…

I suspect a standard 20A 110V circuit can probably handle 2x RTX 6000 Pros. 15A probably can but that requires more research.

During initial setup of the server I am putting together, I found that a machine with 4x Blackwell cards derated to 300W can get by on a single 120V 20A circuit. It's tight but doable. A lot depends on the power supply. I don't think it's a great idea to run 4 high-power GPUs on a single ATX-style PSU, even a beefy 1600W job.

The other questionable part is whether all four cards can temporarily spike at full power during boot, before the wattage limit is applied by the OS. Some accounts say this is possible, and if so it could shut down the party in a hurry. But I didn't see any misbehavior when I tried it.

Re: Was my $48K GPU server worth it?

#142

In the last year, I have bought an M3 Ultra Mac Studio with 512 GB, a Macbook Pro M5 MAX with 128 GB and an RTX 6000 Pro. I have spent around $25k so far, not including electricity. I figured worst case scenario I can sell them in the next year and only take a haircut as opposed to losing my entire investment. In comparison to just spending for tokens, the tokens would have been much cheaper and much much faster. I'v…

I’m not usually one to ask this because learning to do a thing can be fun, but why exactly have you spent 25 thousand dollars on getting an LLM someone else made to answer maths exam questions?

One year ago finetuned local LLMs had a significant edge over ChatGPT or Claude. Look up in YouTube all the DIY videos testing LLMs on their own machines with different setups.

Remember: one year showed up to be a gigantic leap in regards to quality of results and innovation in the AI space. Agents weren't really a thing and vibe coding wasn't even invented as a term because the top notch tools at the time were lousy, with lovable being the frontrunner with its - in my view - sorry Tailwind recombination tool shaming AI to do the work.

Then fall hit 2025 hit us, new year's eve and suddenly there was such a massive surge of innovation and competition with ChatGPT Codex suddenly showing up.

Remember: one year ago many now commonly used tools weren't yet available like Nano Banana or Codex.

"The 25k are so vast" - Yes, and no. For example, if the machine is bought for business usage I can deduct the costs from taxes. This roughly amount for 50% of the financial burden.

So I jokingly use to say, that I pay only half the price for my Apple business machines. And yes, I am strict in this regard. Business means business. No private emails etc. nothing on my company computers.

Maybe there are other options as well to reduce the financial expenses the dude mentions, but it doesn't seem so.

I would also go for leasing, this way already the monthly payments can be deduced and I don't need to buy and maybe resell the machine.

Apple is a luxury good. Without business usage or at least partly using it for business as well as private (mixed usage in tax reports) I wouldn't buy the devices or think twice.

Apple under Cook evolved into a Gucci like luxury brand, that is more and more a rip off than quality delivered, especially considering the latest OS updates for Mac, iOS and iPad. Apple is a mess, following Microsoft Windows' footsteps happily, because the CEO is as has been correctly assessed, no product guy.

But I stop with my rant here.

Always try to use tax deduction as leverage for your computer expenses. Every citizen should invest in basic knowledge about that.

Even a 10-20% professional usage for work (mixed usage) gives you a noticeable advantage over normal pay.

Re: Was my $48K GPU server worth it?

#143

In the last year, I have bought an M3 Ultra Mac Studio with 512 GB, a Macbook Pro M5 MAX with 128 GB and an RTX 6000 Pro. I have spent around $25k so far, not including electricity. I figured worst case scenario I can sell them in the next year and only take a haircut as opposed to losing my entire investment. In comparison to just spending for tokens, the tokens would have been much cheaper and much much faster. I'v…

[deleted]

Re: Was my $48K GPU server worth it?

#144
post #47

Earlier quoted context omitted.

Shallow take: They made an LLM that uses fewer emdashes. Cynical take: They made an LLM that can bypass existing AI slop detectors. Realistic take: They found a research problem they found interesting, dumped a bunch of capital and sweat equity into and (claimed to have, at least) found a solution. Neat!

Or they just have lots of money and a hobby. Someone else might blow $48K to get an old Cessna and go have fun flying around. Not everything needs to have a purpose.

A self-described "broke grad student [who has] been saving up for this for years."

Risking their own money and time instead of leveraging a PowerPoint to hire other peoples' labor with other peoples' money. I can respect that.

Re: Was my $48K GPU server worth it?

#145

I did the math at least on a Macbook pro, and for inference it's definitely not worth it. - https://www.williamangel.net/blog/2026/05/17/offline-llm-ene... - Discussion: https://news.ycombinator.com/item?id=48168198

Except this math is 10x too high (unless accelerated depreciation is all of it) - a million tokens at 28 tokens/sec and 75W and 20c/kwh should cost $0.15 not $1.50. (And less with MTP.)

Re: Was my $48K GPU server worth it?

#146
post #45

Earlier quoted context omitted.

Cloud servers have cheaper electricity, the scale of industrial-level cooling, no issues for you (as a user) with hardware failure (ie you just use a different server; it's not your problem) and can amortize their cost by running 24x7. I've seen H100 computer hours for as little as $2. As the author notes, there are also electrical/wiring issues that cap how much compute gear you can run in a space not designed for i…

I suspect a standard 20A 110V circuit can probably handle 2x RTX 6000 Pros. 15A probably can but that requires more research. During initial setup of the server I am putting together, I found that a machine with 4x Blackwell cards derated to 300W can get by on a single 120V 20A circuit. It's tight but doable. A lot depends on the power supply. I don't think it's a great idea to run 4 high-power GPUs on a single ATX-s…

My earlier research suggests NVIDIA does not actually cap spikes, it caps the average over short periods of time. So setting the power limit is no guarantee.

Re: Was my $48K GPU server worth it?

#147

In the last year, I have bought an M3 Ultra Mac Studio with 512 GB, a Macbook Pro M5 MAX with 128 GB and an RTX 6000 Pro. I have spent around $25k so far, not including electricity. I figured worst case scenario I can sell them in the next year and only take a haircut as opposed to losing my entire investment. In comparison to just spending for tokens, the tokens would have been much cheaper and much much faster. I'v…

If you're in a decent sized city, you should be able to find a local buyer on Craigslist or FB Marketplace... Beyond that, for higher value, smaller items like your M3 Ultra, I would talk to your local police department and/or library to see if you can do the exchange there. Larger libraries usually have a police officer on site or nearby, and the PD office near you may also provide a "safe" exchange location... I'd bring a monitor/keyboard/mouse so you can demonstrate the system working properly.

YMMV but between your nearest PD office and Library, you should be able to use one or the other for your exchange of goods/money. The biggest thing I've sold is a mid-range video card during late covid (I managed to get a better one via newegg shuffle) so I sold the old one (RX 5700XT -> RTX 2080) to make up the difference a bit. I just did the exchange at the Starbucks near me for that.

Re: Was my $48K GPU server worth it?

#148

In the last year, I have bought an M3 Ultra Mac Studio with 512 GB, a Macbook Pro M5 MAX with 128 GB and an RTX 6000 Pro. I have spent around $25k so far, not including electricity. I figured worst case scenario I can sell them in the next year and only take a haircut as opposed to losing my entire investment. In comparison to just spending for tokens, the tokens would have been much cheaper and much much faster. I'v…

I got an RTX 6000 pro too. I like running locally, I've learned a lot more than if I had used an API and there's less worry about overspending tokens. I accidentally spent $100 on claude api in like 2 days because I didn't know what I was doing.

The problem is that while one these gpus is a huge improvement over a laptop or a single 3090, you very quickly wish you had more. I would buy a second one, but I did the math and realized that with the current crop of models, 2 Blackwells doesn't buy me any new capability that I didn't have with one. So I would need a 3rd one. And when I buy a 3rd one I will feel like I want to running a higher quant, so then I will want a 4th.

Re: Was my $48K GPU server worth it?

#149

In the last year, I have bought an M3 Ultra Mac Studio with 512 GB, a Macbook Pro M5 MAX with 128 GB and an RTX 6000 Pro. I have spent around $25k so far, not including electricity. I figured worst case scenario I can sell them in the next year and only take a haircut as opposed to losing my entire investment. In comparison to just spending for tokens, the tokens would have been much cheaper and much much faster. I'v…

Running LLMs on Macs is still terribly slow. They simply lack the optimizations other platforms have. An RTX 6000 pro Blackwell is a pretty good card

M5 pro 48GB should be good and future proof

Re: Was my $48K GPU server worth it?

#150
I administer a simple AI server in the office, which just uses a single RTX 5090 but is able to serve ~80 people throughout the day. I'm impressed by Qwen3.6-27b's capabilities in agentic coding/tasks so far. Devs say it's not much different from Sonnet 4.6 on many tasks (sometimes it even outperformed it), 40-60 tok/sec, up to 260k context. The server cost about $10k with all the bells and whistles.

I spent a lot of time researching/adding/benchmarking many custom modifications to the software stack and its settings to make the server optimally handle the load with just 1 RTX 5090 without losing quality, but it's still not enough, and the wait times in the queue are getting longer. We're at the limits of the hardware, and I'm out of tricks.

The experiment was kind of a success, and the CTO agrees we should scale it. With our own infra, we could run agents 24/7 on everything. Currently, a lot of use cases for the cloud providers are completely blocked by PII/trade secret concerns (our infosec department doesn't buy the "zero retention" promise), plus you don't have to think about billing/budgets/etc. anymore.

Now I can't decide how to scale it. On one hand, I'd like to run larger models. And we have the budget to buy, say, 8xH200. But in many benchmarks, the larger models that do fit in 8xH200 comfortably and can serve many parallel requests with acceptable speed/quality don't seem to outperform Qwen3.6 that much in agentic coding/tasks to justify the price.

So another option is just to buy a bunch of RTX 6000s and scale horizontally instead: run a copy of a midrange LLM like Qwen3.6 on each GPU. It's cheaper and easier to scale/replace, but then we'll run into problems running larger models in the future if we have to, because of no NVLink support (say, if Alibaba & Co. stop releasing ~30b models and/or ~30b models start falling behind 400b+ models considerably)

Does anyone here have experience running large models in a multi-GPU setup with several RTX 6000s in a high-concurrency regime and with large context lengths? (something like Deepseek 4 Flash, Minimax 2.7 etc.)

Post reply on HN