Live data from Hacker News

Was my $48K GPU server worth it?

rosmine.ai

321–330 of 480 posts

Re: Was my $48K GPU server worth it?

#321

Earlier quoted context omitted.

On top of that, AI providers are also eating a big loss on the service.

Supposedly Anthropic just reported that they’re operationally profitable. So maybe not?

"operationally" implies that capex (which I would assume includes datacenters, gpus, and r&d) is not in. So the big news is that they can now pay for electricity and sysadmin.

Re: Was my $48K GPU server worth it?

#322

"The point of buying the server wasn’t to save money, it was to build something cool." In the end, this is always the real answer - one that I'm sure we can all agree is the 'correct' one too.

I'm sure thr plan was to build a holodeck all the time..

Re: Was my $48K GPU server worth it?

#323
post #306

Earlier quoted context omitted.

Are they? I only ever see unsubstantiated claims for this whereas I see many justifications that interference is comfortably profitable in isolation.

Its basic math, go calculate max sessions for a certain tps on any hardware. Session# * tps * 86400 (secs in a day) * 30 days. You'll realize real quick its not profitible. You cant just say things you don't like to hear are unsubstantiated without verifying. Not to mention, subscriptions.. $2mm in GPUs being given out for 5 hrs a day at a cost of $200 a month. I could easily say that everyone who says its profitible…

>Its basic math

Yes, once you have modeled the problem correctly and you know all the input parameters. This is not that: Session# * tps * 86400 (secs in a day) * 30 days.

I don't think there is enough public information to check Anthropic's claims regarding inference profitability. It depends not just on unknown technical factors but also on agreements they have with other companies.

Re: Was my $48K GPU server worth it?

#324

Earlier quoted context omitted.

You got numbers? Because it seems perfectly possible to me. OpenAI and Anthropic’s marginal cost for inference is certainly far less than their API pricing.

How can you say that with such certainty? You have no idea what it costs to run a 10T parameter model at extremely high concurrency. These 1T param models running at <$3.00 per 1mm are certainly not profitable.

Because I’ve looked at what it would cost my company to self-host a SOTA sized model. For us it wasn’t worth it because the hardware is all bought up by frontier labs and we can’t get any supply. But if we could, at the prices they’re paying, it would pay for itself in 10-ish months. I assume further that they have economies of scale on top of what I was estimating.

Re: Was my $48K GPU server worth it?

#325

Earlier quoted context omitted.

oh young grasshopper, I see you dont know that money launderers love the ebay hype cycle. Its REALLY common on high dollar hot items to have phantom transactions where parties are on both sides of the transaction to clean illicit money. The high price tag and high volume amount of transactions hides the illicit signal. I have tried to buy a few of these mac studios only to have the transaction cancelled because I was…

Bizarre. They don't care that eBay takes 14%?

14% seems like a pretty low fee to clean drug money, if we're being honest

Re: Was my $48K GPU server worth it?

#326
Any kind of fixed capacity usage model seems to be a dead end. Paying per token might seem like an exploitative arrangement at first glance, but it's a luxury if you are experimenting or deploying greenfield.

Provisioned capacity is a really high end thing. I feel like you'd need to be spending more than $1000/day on tokens for this model to make any sense. You lose a lot of flexibility once you start dumping capital into specific pieces of hardware. Maybe start by renting the GPU server for a few days...

Re: Was my $48K GPU server worth it?

#327

Earlier quoted context omitted.

You spent 50k for plex hosting? Why so expensive?

Half a petabyte of RAID6 is the biggest line item, then the redundant 40gb networking and compute follow closely. I have a lot… too much even?

does one really need 40gb networking to stream bluerays?

Re: Was my $48K GPU server worth it?

#328
post #306

Earlier quoted context omitted.

Are they? I only ever see unsubstantiated claims for this whereas I see many justifications that interference is comfortably profitable in isolation.

Its basic math, go calculate max sessions for a certain tps on any hardware. Session# * tps * 86400 (secs in a day) * 30 days. You'll realize real quick its not profitible. You cant just say things you don't like to hear are unsubstantiated without verifying. Not to mention, subscriptions.. $2mm in GPUs being given out for 5 hrs a day at a cost of $200 a month. I could easily say that everyone who says its profitible…

We should specify which subscription plan we are talking about. You seem to be talking about the Anthropic Claude Max plan. I think it's consensus that these flat rate type of subscriptions are loss leaders, as they come with restrictions how you can use the API via T&C, namely only with Claude Code et al. They are meant to hook developers into their products.

Shouldn't we compare the API pricing, where we pay per token? The whole point of local inference is that we don't have any restrictions regarding product use or time limits, so it would only be fair if we compare it to a plan that offers the same. And even that is only a first approximation, because the commercial models are usually much more capable than the open weight models.

Re: Was my $48K GPU server worth it?

#329

> The mentality shift of renting vs. owning the gpus is huge. When renting, each experiment costs money and I had to ask myself is it worth it. When owning, it feels like not running experiments is costing me money. I feel like there is some very deep generalizable wisdom buried here.

Also something about subscriptions vs pay-for-usage. I feel the need to use all my weekly tokens or I'm wasting and I bet they would never get this kind of usage out of me if AI ended up being same price per token.
Post reply on HN