Live data from Hacker News

Show HN: San Francisco Compute – 512 H100s at <$2/hr for research and startups

sfcompute.org

11–20 of 189 posts

Re: Show HN: San Francisco Compute – 512 H100s at <$2/hr for research and startups

#11

I hope you succeed. TPU research cloud (TRC) tried this in 2019. It was how I got my start. In 2023 you can barely get a single TPU for more than an hour. Back then you could get literally hundreds, with an s. I believed in TRC. I thought they’d solve it by scaling, and building a whole continent of TPUs. But in the end, TPU time was cut short in favor of internal researchers — some researchers being more equal than…

> In 2023 you can barely get a single TPU for more than an hour. Um. Can't you order them from coral.ai and put them in an NVMe slot? Or are the cloud TPUs more powerful?

TPU pod is not sold by google, edge tpu is different

Re: Show HN: San Francisco Compute – 512 H100s at <$2/hr for research and startups

#12

Earlier quoted context omitted.

> In 2023 you can barely get a single TPU for more than an hour. Um. Can't you order them from coral.ai and put them in an NVMe slot? Or are the cloud TPUs more powerful?

TPU pod is not sold by google, edge tpu is different

So the cloud TPUs are more powerful...? Or what are you saying?

Re: Show HN: San Francisco Compute – 512 H100s at <$2/hr for research and startups

#14

Earlier quoted context omitted.

TPU pod is not sold by google, edge tpu is different

So the cloud TPUs are more powerful...? Or what are you saying?

Yeah, it’s a silly branding thing.

One TPU (not even a pod, just a regular old TPUv2) has 96 CPU cores with 1.4TB of RAM, and that’s not even counting their hardware acceleration. I’d love to buy one.

Re: Show HN: San Francisco Compute – 512 H100s at <$2/hr for research and startups

#19

How does this compare to https://lambdalabs.com/ ?

Very similar price, but from what I gather very different model. One important difference might be if you regularly run short-ish training runs over many GPUs. Lambdalabs might not have 256 instances to give you right now. With OP you are basically buying the right to put jobs in the job queue for their 512 GPU cluster, so running a job that needs 256 GPUs isn't an issue (though you might wait behind someone running a 512 GPU job).

No idea how capacity at lambdalabs actually looks like though. Does anyone have insight how easy it is to spin up more than 2-3 instances up there?

Re: Show HN: San Francisco Compute – 512 H100s at <$2/hr for research and startups

#20

How does this compare to https://lambdalabs.com/ ?

Ah, we're running a medium amount of compute at zero-margin. The point is not to go sell the Fortune 500, but to make sure a grad student can spend a $50k grant.

Right now, it's pretty easy to get a few A/H100s (Lambda is great for this), but very hard to get more than 24 at a reasonable price ($~2 an hour). One often needs to put up a 6+ month commitment, even when they may only want to run their H100s for an 8 hour training run.

It's the right business decision for GPU brokers to do long term reservations and so on, and we might do so too if we were in their shoes. But we're not in their shoes and have a very different goal: arm the rebels! Let someone who isn't BigCorp train a model!

Post reply on HN