Live data from Hacker News

Show HN: San Francisco Compute – 512 H100s at <$2/hr for research and startups

sfcompute.org

141–150 of 189 posts

Re: Show HN: San Francisco Compute – 512 H100s at <$2/hr for research and startups

#141
post #99

I hope you succeed. TPU research cloud (TRC) tried this in 2019. It was how I got my start. In 2023 you can barely get a single TPU for more than an hour. Back then you could get literally hundreds, with an s. I believed in TRC. I thought they’d solve it by scaling, and building a whole continent of TPUs. But in the end, TPU time was cut short in favor of internal researchers — some researchers being more equal than…

Actually, the TPU Research Cloud program is still going strong! We've expanded the compute pool significantly to include Cloud TPU v4 Pod slices, and larger projects still use hundreds of chips at a time. (TRC capacity has not been reclaimed for internal use.) Check out this list of recent TRC-supported publications: https://sites.research.google/trc/publications/ Demand for Cloud TPUs is definitely intense, so if yo…

Main problem with the TPU Research Cloud is you get dragged down a LOT by the buggy TPU API-- not just the Google Cloud API being awful but the Tensorflow/Jax/Pytorch support also being awful too. You also basically must use Google Cloud Storage, which is also slow and can be really expensive getting anything into / out-of.

The Googlers maintaining the TPU Github repo also just basically don't care about your PR unless it's somehow gonna help them in their own perf review.

In contrast with a GPU-based grid, you can not only run the latest & greatest out-of-the-box but also do a lot of local testing that saves tons of time.

Finally, the OP here appears to be offering real customer engagement, which is totally absent from my own GCloud experiences across several companies.

Re: Show HN: San Francisco Compute – 512 H100s at <$2/hr for research and startups

#142
post #123

Earlier quoted context omitted.

A few quick comments: > But it’s important for hobbyists and tinkerers to be able to participate in the AI ecosystem Totally agree! This was a big part of my original motivation for creating the TPU Research Cloud program. People sometimes assume that e.g. an academic affiliation is required to participate, but that isn't true; we want the program to be as open as possible. We should find a better way to highlight th…

It’s not euw4a. It’s everywhere. The allocation algorithm across the board kills off TPUs after no more than a couple hours. usc1f, usc1a, usc1c, euw4a; it makes no difference. It would be funny if someone set gpt-2-15b-poetry (our project) in some special way to prevent us from making TPUs that ever last more than a few hours, but from what I’ve heard from other people, this isn’t the case. That’s what I mean about…

It sounds like you're primarily using preemptible TPU quota, which doesn't come with any availability or uptime expectations at all.

By default, the TRC program grants both on-demand quota and preemptible quota. If you are able to create a TPU VM with your on-demand quota, it should last quite a bit longer than a few hours. (There are situations in which on-demand TRC TPU VMs can be interrupted, but these ought to be rare.) If your on-demand TPU VMs are being interrupted frequently, please email TRC support and provide the names of the TPU hosts that were interrupted so folks can try to help.

When there is very high demand for Cloud TPUs, it's certainly possible for preemptible TPU VMs to be interrupted frequently. It would be an interesting engineering project to make a very robust training system that could make progress even with low TPU VM uptime, and I hope someone does it! Until then, though, you should have a better experience with on-demand resources when you're able to create them. Reserved capacity is even better since it provides an expectation of both availability and uptime.

Re: Show HN: San Francisco Compute – 512 H100s at <$2/hr for research and startups

#144
post #120

Earlier quoted context omitted.

> You don’t need to take my word for it. Here’s some unfiltered DMs on the subject: https://imgur.com/a/6vqvzXs > Notice how their optimism dries up, and not because I was telling them how bad TRC has become. It’s because their TPUs kept dying. Unless I'm misreading this they sound pretty happy and you sound pessimistic? Their last substantial comment was "I'm sure Zak could hook you up with something better"?

TRC is supposed to be the “something better”. This insider TPU stuff is for the birds. If TRC can only offer 4 hours with no preemptions, that’s fine, but they need to be up front about that. Saying that TPUs preempt every 24 hours and then killing them off after 45 minutes is… not very productive. As for their comments, the third screenshot is the key; they’re agreeing that the situation is bad. They’re a friend, an…

As mentioned in another comment, it sounds like you're using preemptible TRC TPU quota. If you use on-demand TRC TPU quota instead, that should improve your uptime substantially.

Re: Show HN: San Francisco Compute – 512 H100s at <$2/hr for research and startups

#145
post #99

Earlier quoted context omitted.

Actually, the TPU Research Cloud program is still going strong! We've expanded the compute pool significantly to include Cloud TPU v4 Pod slices, and larger projects still use hundreds of chips at a time. (TRC capacity has not been reclaimed for internal use.) Check out this list of recent TRC-supported publications: https://sites.research.google/trc/publications/ Demand for Cloud TPUs is definitely intense, so if yo…

Main problem with the TPU Research Cloud is you get dragged down a LOT by the buggy TPU API-- not just the Google Cloud API being awful but the Tensorflow/Jax/Pytorch support also being awful too. You also basically must use Google Cloud Storage, which is also slow and can be really expensive getting anything into / out-of. The Googlers maintaining the TPU Github repo also just basically don't care about your PR unle…

Could you share a few technical details about the issues you've encountered with TF / JAX / PyTorch on Cloud TPUs? The overall Cloud TPU user experience improved a whole lot when we enabled direct access to TPU VMs, and I believe the newer JAX and PyTorch integrations are improving very rapidly. I'd love to know which issues are currently causing the most friction.

Re: Show HN: San Francisco Compute – 512 H100s at <$2/hr for research and startups

#146
post #20

Earlier quoted context omitted.

Ah, we're running a medium amount of compute at zero-margin. The point is not to go sell the Fortune 500, but to make sure a grad student can spend a $50k grant. Right now, it's pretty easy to get a few A/H100s (Lambda is great for this), but very hard to get more than 24 at a reasonable price ($~2 an hour). One often needs to put up a 6+ month commitment, even when they may only want to run their H100s for an 8 hour…

So what happens when some big bucks VC backed closed source LLM company buys all your compute inventory for the next 5 years? This is not that unlikely. Lambda Labs a little while back was completely sold out of all compute inventory.

Presumably them buy more gpus

Re: Show HN: San Francisco Compute – 512 H100s at <$2/hr for research and startups

#147
post #132

Nat Friedman and Daniel Gross setup a 2,512 H100 cluster [1] for their startups, with a very similar “shared” model. Might be interesting to connect with them. [1] https://andromedacluster.com/

Nat & Daniel’s cluster is great, and we fully recommend startups seek out this option as well. Nat & Daniel are some of the best investors one can have

Re: Show HN: San Francisco Compute – 512 H100s at <$2/hr for research and startups

#148
Honest question I don’t know how to consider: are we further along or behind with AI given crypto’s use of GPUs? Has the same cards bought for mining furthered AI, or maybe that demand lead to more research into GPUs and what they can do - or would we be further along if we weren’t wasting these cards on mining?

Re: Show HN: San Francisco Compute – 512 H100s at <$2/hr for research and startups

#149

Honest question I don’t know how to consider: are we further along or behind with AI given crypto’s use of GPUs? Has the same cards bought for mining furthered AI, or maybe that demand lead to more research into GPUs and what they can do - or would we be further along if we weren’t wasting these cards on mining?

Ethereum's (thrice delayed) move to PoS put a glut of GPUs on the market, just in time for the AI boom to swallow them back up, so I think it ended up okay. NVDA certainly had a great few days in the market thanks to AI though.

Re: Show HN: San Francisco Compute – 512 H100s at <$2/hr for research and startups

#150

Earlier quoted context omitted.

Huh, this doesn't seem right. Based on #s you seem to be referring to pods but even then I'm not familiar with such a configuration existing. A single TPUv2 chip has 1 core and 8gb of memory. A single device comes in the v2-8 configuration with 8 cores and 64gb of memory. Pod variants come in v2-32 to v2-512 configurations.

A single TPUv2 host has 8 TPU cores with 64GB of total HBM (8GB per core), but like GPUs, TPUs can't directly access a network, so the host also needs CPUs and standard RAM to send data to them. They are fast, and the host has to be fast enough to keep them fed with data, so the host is pretty beefy. But FWIW, a TPUv2 host has somewhere around 330GB of RAM, not 1.4TB.

Thanks for clarifying, I misinterpreted the commenter as referring to the accelerator as the conversation was about TPU availability for purchase.

I know just enough about the architecture to facilitate using TPUs for research training runs but I'm not sure what's so special about the host?

Sure it's beefy but there are much beefier servers readily available.

Post reply on HN