Live data from Hacker News

Show HN: San Francisco Compute – 512 H100s at <$2/hr for research and startups

sfcompute.org

151–160 of 189 posts

Re: Show HN: San Francisco Compute – 512 H100s at <$2/hr for research and startups

#151
post #120

Earlier quoted context omitted.

> You don’t need to take my word for it. Here’s some unfiltered DMs on the subject: https://imgur.com/a/6vqvzXs > Notice how their optimism dries up, and not because I was telling them how bad TRC has become. It’s because their TPUs kept dying. Unless I'm misreading this they sound pretty happy and you sound pessimistic? Their last substantial comment was "I'm sure Zak could hook you up with something better"?

TRC is supposed to be the “something better”. This insider TPU stuff is for the birds. If TRC can only offer 4 hours with no preemptions, that’s fine, but they need to be up front about that. Saying that TPUs preempt every 24 hours and then killing them off after 45 minutes is… not very productive. As for their comments, the third screenshot is the key; they’re agreeing that the situation is bad. They’re a friend, an…

This is totally fascinating.

Frankly, it sounds to me like they're having severe yield+reliability problems with the TPUv4s that aren't getting caught by wafer-level testing, and have binned the flakiest ones for use by outsiders.

A lot of yield issues show up as spontaneous resets/crashes.

Re: Show HN: San Francisco Compute – 512 H100s at <$2/hr for research and startups

#152

Earlier quoted context omitted.

TRC is supposed to be the “something better”. This insider TPU stuff is for the birds. If TRC can only offer 4 hours with no preemptions, that’s fine, but they need to be up front about that. Saying that TPUs preempt every 24 hours and then killing them off after 45 minutes is… not very productive. As for their comments, the third screenshot is the key; they’re agreeing that the situation is bad. They’re a friend, an…

This is totally fascinating. Frankly, it sounds to me like they're having severe yield+reliability problems with the TPUv4s that aren't getting caught by wafer-level testing, and have binned the flakiest ones for use by outsiders. A lot of yield issues show up as spontaneous resets/crashes.

It's more likely Google preempting researcher who are on a preemptable research grant, and it is happening a lot more often because there are more paying customers.

Re: Show HN: San Francisco Compute – 512 H100s at <$2/hr for research and startups

#153
post #152

Earlier quoted context omitted.

This is totally fascinating. Frankly, it sounds to me like they're having severe yield+reliability problems with the TPUv4s that aren't getting caught by wafer-level testing, and have binned the flakiest ones for use by outsiders. A lot of yield issues show up as spontaneous resets/crashes.

It's more likely Google preempting researcher who are on a preemptable research grant, and it is happening a lot more often because there are more paying customers.

"Preemptable money" sounds like the kind of bullshit I would use to cover up failed chips. And yes, I am a VLSI engineer.

Re: Show HN: San Francisco Compute – 512 H100s at <$2/hr for research and startups

#154
post #130
post #70

> Rather than each of K startups individually buying clusters of N gpus, together we buy a cluster with NK gpus... Then we set up a job scheduler to allocate compute In theory, this sounds almost identical to the business model behind AWS, Azure, and other cloud providers. "Instead of everyone buying a fixed amount of hardware for individual use, we'll buy a massive pool of hardware that people can time-share." Outsi…

Couple things, mostly pricing and availability: 1) Margins. Public cloud investors expect a certain margin profile. They can’t compete with Lambda/Fluidstack’s margins. 2) To an extent also big clouds have worse networking for LLM training. I believe only Azure has infiniband. Oracle is 3200 Gbps but not infiniband, same for AWS I believe. GCP not sure but their A100 networking speeds were only 100 Gbps I believe rat…

Low margins and “will this thing still be around in 2 years” are negatively correlated.

Where’s the capital for upgrades, repairs, and replacements coming from?

Re: Show HN: San Francisco Compute – 512 H100s at <$2/hr for research and startups

#155

Earlier quoted context omitted.

FWIW I tired to look up some numbers, i found California "industrial" electricity at $0.18/Kwh https://www.eia.gov/electricity/monthly/epm_table_grapher.ph... and H100s using 300-700w https://www.nvidia.com/en-us/data-center/h100/ which implies a worst case marginal cost of .18*.7 = $.126 / gpu / hour. Looks like Montana is cheapest at ~$.05 / kwh which would bring that down to $.035. So there may be about a $0.09 Ca…

Retail residential power in the city of Santa Clara is $0.15/KwH, I'm sure commercial could be less. Especially if you throw some solar panels on the roof. The most expensive part would be the land, but honestly there is some pretty cheap land outside the cities.

For reference, I'm in SF and paid PGE $0.50938/KWh during peak hours, residential, last bill.

Re: Show HN: San Francisco Compute – 512 H100s at <$2/hr for research and startups

#157
post #130
post #70

> Rather than each of K startups individually buying clusters of N gpus, together we buy a cluster with NK gpus... Then we set up a job scheduler to allocate compute In theory, this sounds almost identical to the business model behind AWS, Azure, and other cloud providers. "Instead of everyone buying a fixed amount of hardware for individual use, we'll buy a massive pool of hardware that people can time-share." Outsi…

Couple things, mostly pricing and availability: 1) Margins. Public cloud investors expect a certain margin profile. They can’t compete with Lambda/Fluidstack’s margins. 2) To an extent also big clouds have worse networking for LLM training. I believe only Azure has infiniband. Oracle is 3200 Gbps but not infiniband, same for AWS I believe. GCP not sure but their A100 networking speeds were only 100 Gbps I believe rat…

What is your differentiator from Lambda? That you are smaller and in a single DC?

Sincere question.

Re: Show HN: San Francisco Compute – 512 H100s at <$2/hr for research and startups

#158
Please take this question without prejudice.

Is it accurate to say you’re willing to go into ~20,000,000 USD debt to sell discounted computer-as-a-service to researchers/startups, but unwilling to go into debt to sponsor the undergraduate degrees of ~100-500 students at top-tier schools? (40k - 200k USD per degree)

Or, you know, build and fund a small public school/library or two for ~5 years?

Re: Show HN: San Francisco Compute – 512 H100s at <$2/hr for research and startups

#159

I hope you succeed. TPU research cloud (TRC) tried this in 2019. It was how I got my start. In 2023 you can barely get a single TPU for more than an hour. Back then you could get literally hundreds, with an s. I believed in TRC. I thought they’d solve it by scaling, and building a whole continent of TPUs. But in the end, TPU time was cut short in favor of internal researchers — some researchers being more equal than…

> In 2023 you can barely get a single TPU for more than an hour

Oh come on, colab gives TPU access in the free tier for a whole half day. No need to exaggerate the shortage

Re: Show HN: San Francisco Compute – 512 H100s at <$2/hr for research and startups

#160
post #134

I know AWS/GCP/Azure have overhead and I understand why so many companies choose to go bare metal on their ops. I personally rarely think it's worth the time and effort, but I get that with scale saving can be substantial. But for AI training? If the public cloud isn't competitive even for bursty AI training, their margins are much higher than I anticipated. OP mentions 10-20x cost reduction? Compared to what? AWS?

AWS offers p5.48xlarge which is 8xH100 for $98.32, so 12.29$ per hour per H100 - ~6x the price.
Post reply on HN