Live data from Hacker News

Nobody knows what a used GPU cluster is worth

ciphertalk.substack.com

281–285 of 285 posts

Re: Nobody knows what a used GPU cluster is worth

#281

Earlier quoted context omitted.

"Sounds like a plane taking off." Helicopter. I live in Yeovil, Somerset, UK - there's a helicopter factory just down the road. I had a IBM "AS/400" or whatever they are called now in our computer room rack for a customer and it made nearly as much noise as everything else put together. It was clearly tuned for start up noise to impress because they would fire up in sequence, rise to a crescendo and then slow down in…

That's called staggered spin-up. One does that for sets of HDDs, too. Anyway, what is it good for? To not overload the PSU for one, because spin-up needs more energy, until it has reached it's intended range of RPM. (for fans & HDDs) Another reason to spin-up high initially, to fall down to slower when up , is to overcome friction in the bearings, to get stuck things unstuck, and shake dust off. (for fans)

Tell that to my switches and PC servers. They all spin up their fans in one go as soon as power good is received and the initial BIOS thingie has stubbed out its fag on the floor.

The PSUs are painting their nails when the fans go off initially. The GPUs, RAID controllers and co are still waking up.

A further counter example is an Equallogic PS6500. That has 48 3.5" spindles in it and three PSUs plus a few fans and two controllers. The whole thing boots in a couple of minutes.

Re: Nobody knows what a used GPU cluster is worth

#282
post #275

Earlier quoted context omitted.

Because llm inference is not the only workload a GPU can do and custom silicon cost $10b a chip.

AI optimised cards is the biggest cashcow for nvidia. They are definitely doing custom silicon. And Google et al have cards that don't even pretend to be able to do graphics.

There is more than one workload in AI. Inference for llms is memory constrained on even a single card. Training for llms is memory constrained on the level of racks. In both cases you hardly ever see more than 40% of advertised flops used.

Re: Nobody knows what a used GPU cluster is worth

#283
post #98

Earlier quoted context omitted.

It was true of crypto GPUs too, although mostly people picking them up for gaming. Always seems high to me too but if you can get any guarantee of them not being on fire when they were pulled the bathtub curve keeps you pretty safe, thermal limits are limits for a reason.

crypto gpus didnt run at 100% power draw so the strain wasn't that massive

we over clocked and under volted... the goal was to lower power and run the clocks as fast as possible.

Re: Nobody knows what a used GPU cluster is worth

#284
post #275

Earlier quoted context omitted.

AI optimised cards is the biggest cashcow for nvidia. They are definitely doing custom silicon. And Google et al have cards that don't even pretend to be able to do graphics.

There is more than one workload in AI. Inference for llms is memory constrained on even a single card. Training for llms is memory constrained on the level of racks. In both cases you hardly ever see more than 40% of advertised flops used.

So the question is: why are they designing AI cards so unbalanced?

Re: Nobody knows what a used GPU cluster is worth

#285
post #64

I like Meg a lot a human, but Meg is all doom and gloom. Every single post she makes is about how GPUs fail [0] and now she's onto how financing is a big thing just waiting to crash and "nobody knows what a used GPU cluster is worth"... Actually, we do, people offer them to me all the time. A used box of MI300x is $257k. "There is no GPU futures market"... actually there are a few of them that people have pitched to…

Aren't the AI Datacenters for the Digital Control Grid, so govt contracts?

This wasn't sarcasm. Gov't contracts are where a lot of people make a lot of money. And Palantir and ICE are definitely a thing. I know that sounds purely political but these are real things. ICE depends on Palantir a lot. And Palantir uses Data Center resources a lot.

Why is the Jensen the CEO of NVidia lending out huge sums of money for AI companies to buy their chips for datacenters? Part of the reason is the security of government contracts ala the Digital Control Grid.

Post reply on HN