Having hosted infrastructure in CA at multiple colos. I would advise you to host it elsewhere if you can, cost of power, other infrastructure is much higher in CA than AZ or NV.
Show HN: San Francisco Compute – 512 H100s at <$2/hr for research and startups
101–110 of 189 posts
Re: Show HN: San Francisco Compute – 512 H100s at <$2/hr for research and startups
#102Earlier quoted context omitted.
AWS and Azure would slit their own throats before they created a way for their customers to pool instances to save money. They want to do that themselves, and keep the customer relationship and the profits, instead of giving them to a middleman or the customer.
It’s just corporate profits combined with market forces, not a some sort of malicious conspiracy. You can rent a 2-socket AMD server with 120 available cores and RDMA for something like 50c to $2 per hour. That’s just barely above the cost of the electricity and cooling! What do you want, free compute just handed to you out of the goodness of their hearts? There is incredible demand for high-end GPUs right now, and m…
Re: Show HN: San Francisco Compute – 512 H100s at <$2/hr for research and startups
#103Who is funding this?
Cause if it’s VC then it’s going to have the same fate as everything else after 5-7 years.
I hope y’all have as innovative of a business model. You’ll need it if you want to do what you’re doing now for more than a few years
Re: Show HN: San Francisco Compute – 512 H100s at <$2/hr for research and startups
#104How does this compare to https://lambdalabs.com/ ?
Ah, we're running a medium amount of compute at zero-margin. The point is not to go sell the Fortune 500, but to make sure a grad student can spend a $50k grant. Right now, it's pretty easy to get a few A/H100s (Lambda is great for this), but very hard to get more than 24 at a reasonable price ($~2 an hour). One often needs to put up a 6+ month commitment, even when they may only want to run their H100s for an 8 hour…
Re: Show HN: San Francisco Compute – 512 H100s at <$2/hr for research and startups
#105Earlier quoted context omitted.
AWS and Azure would slit their own throats before they created a way for their customers to pool instances to save money. They want to do that themselves, and keep the customer relationship and the profits, instead of giving them to a middleman or the customer.
It’s just corporate profits combined with market forces, not a some sort of malicious conspiracy. You can rent a 2-socket AMD server with 120 available cores and RDMA for something like 50c to $2 per hour. That’s just barely above the cost of the electricity and cooling! What do you want, free compute just handed to you out of the goodness of their hearts? There is incredible demand for high-end GPUs right now, and m…
Re: Show HN: San Francisco Compute – 512 H100s at <$2/hr for research and startups
#106I hope you succeed. TPU research cloud (TRC) tried this in 2019. It was how I got my start. In 2023 you can barely get a single TPU for more than an hour. Back then you could get literally hundreds, with an s. I believed in TRC. I thought they’d solve it by scaling, and building a whole continent of TPUs. But in the end, TPU time was cut short in favor of internal researchers — some researchers being more equal than…
Actually, the TPU Research Cloud program is still going strong! We've expanded the compute pool significantly to include Cloud TPU v4 Pod slices, and larger projects still use hundreds of chips at a time. (TRC capacity has not been reclaimed for internal use.) Check out this list of recent TRC-supported publications: https://sites.research.google/trc/publications/ Demand for Cloud TPUs is definitely intense, so if yo…
This may feel like an anime betrayal, since you basically launched my career as a scientist. But it’s important for hobbyists and tinkerers to be able to participate in the AI ecosystem, especially today. And TRC just does not support them anymore. I tried, many times, over the last year and a half.
You don’t need to take my word for it. Here’s some unfiltered DMs on the subject: https://imgur.com/a/6vqvzXs
Notice how their optimism dries up, and not because I was telling them how bad TRC has become. It’s because their TPUs kept dying.
I held out hope for so long. I thought it was temporary. It ain’t temporary, Zak. And I vividly remember when it happened. Some smart person in google proposed a new allocation algorithm back near the end of 2021, and poof, overnight our ability to make TPUs went from dozens to a handful. It was quite literally overnight; we had monitoring graphs that flatlined. I can probably still dig them up.
I’ve wanted to email you privately about this, but given that I am a small fish in a pond that’s grown exponentially bigger, I don’t think it would’ve made a difference. The difference is in your last paragraph: you allocate reserved instances to those who deserve it, and leave everybody else to fight over 45 minutes of TPU time when it takes 25 minutes just to create and fill your TPU with your research data.
Your non-preemptible TPUs are frankly a lie. I didn’t want to drop the L word, but a TPUv3 in euw4a will literally delete itself — aka preempt — after no more than a couple hours. I tested this over many months. That was some time ago, so maybe things have changed, but I wouldn’t bet on it.
There’s some serious “left hand doesn’t know that right hand detached from its body and migrated south for the winter” energy in the TRC program. I don’t know where it embedded itself, but if you want to elevate any other engineers from software devs to researchers, I urge you to make some big changes.
One last thing. The support staff of TRC is phenomenal. Jonathan Colton has worked more miracles than I can count, along with the rest of his crew. Ultimately he had to send me an email like “by the way, TRC doesn’t delete TPUs. This distinction probably won’t be too relevant, but I wanted to let you know” (paraphrasing). Translation: you took the power away from the people who knew where to put it (Jonathan) and gave it to some really important researchers, probably in Brain or some other division of Google. And the rest is history. So I don’t want to hear that one of the changes is “ok, we’ve punished the support staff” - as far as I can tell, they’ve moved mountains with whatever tools they had available, and I definitely wouldn’t have been able to do any better in their shoes.
Also, hello. Thanks for launching my career. Sorry that I had to leave this here, but my duty is to the open source community. The good news is that you can still recover, if only you’d revert this silly “we’ll slip you some reserved TPUs that don’t kamikaze themselves after 45 minutes if you ask in just the right way” stuff. That wasn’t how the program was in 2019, and I guarantee that I couldn’t have done the work I did then under the current conditions.
Re: Show HN: San Francisco Compute – 512 H100s at <$2/hr for research and startups
#107> Rather than each of K startups individually buying clusters of N gpus, together we buy a cluster with NK gpus... Then we set up a job scheduler to allocate compute In theory, this sounds almost identical to the business model behind AWS, Azure, and other cloud providers. "Instead of everyone buying a fixed amount of hardware for individual use, we'll buy a massive pool of hardware that people can time-share." Outsi…
They are working on this. All the major clouds have initiatives to do short term requests/reservations. It’s just not a feature that has ever been of much use pre-GenAI. How often do you need to request 1000 CPU nodes for 48 hours in a single zone? Secondly, there is a fundamental question of resource sharing here. Even with this project by Evan and AI Grant (the second such cluster created by AI Grant btw), the ques…
I would srgue this has always been a common case for cloud GPU compute
Re: Show HN: San Francisco Compute – 512 H100s at <$2/hr for research and startups
#108The billion dollar question is: Who is funding this? Cause if it’s VC then it’s going to have the same fate as everything else after 5-7 years. I hope y’all have as innovative of a business model. You’ll need it if you want to do what you’re doing now for more than a few years
Not everything has grow to have the appetite of Galactus and swallow a whole planet. Making single digit millions of dollars over a couple of years is still worthwhile, especially if it helps others and moves humanity forwards.
This project isn't ever going to want to try and compete with AWS, so no, it's not a billion dollar question. $20 Million, yeah.
Re: Show HN: San Francisco Compute – 512 H100s at <$2/hr for research and startups
#109Earlier quoted context omitted.
From sentence one of the post, it clearly states that they are VC funders who are doing this for a round of startups they just funded, and they're looking for others to be a part of it.
Oh no, definitely not. We just got a loan. Neither Alex or I are currently VCs, and this has no affiliation with any venture fund. We want to be a customer of the sf compute group too!
If not and you got the loan from a bank, super curious how you were able to get the bank to trust that renting out the GPUs would cover the loan or if some other reasoning convinced them. Assuming you aren’t trying to turn this into a big business, that knowledge might help a lot of other players run similar programs and further democratize SOA GPU access.
Re: Show HN: San Francisco Compute – 512 H100s at <$2/hr for research and startups
#110Earlier quoted context omitted.
> Requirements > Ubuntu 18.04 or newer (required) > Dedicated machines only - the machine shouldn't be doing other stuff while rented well that's certainly not what I expected. ctrl-f "virtual" gives nothing, so it seems they really mean "take over your machine" > Note: you may need to install python2.7 to run the install script. what kind of nonsense is this? Did they write the script in 2001 and just abandon it?
I just skimmed their FAQ at https://vast.ai/faq , and it seems like it could use an update. E.g., it says "Initially we are supporting Ubuntu Linux, more specifically Ubuntu 16.04 LTS.". That version of Ubuntu has been end-of-life'd for several years, and when I just tried vast.ai out, it seemed to be using Ubuntu 20.04. There were also a couple of words with letters missing (probably trivial typos) that could be fou…
It's unique in that you can set your own prices, it's a true spot marketplace.. I've grabbed 2x3090 for $0.02/hr before.
Probably no good for training (can be interrupted any time with zero warning ssh just drops and that's it) but for my inference usecases it lets me spot heavy compute for pennies.