Live data from Hacker News

Show HN: San Francisco Compute – 512 H100s at <$2/hr for research and startups

sfcompute.org

121–130 of 189 posts

Re: Show HN: San Francisco Compute – 512 H100s at <$2/hr for research and startups

#121
post #95
post #67

Earlier quoted context omitted.

>Optimism is (almost) always required in order to accomplish anything of significance. Those who lose it, aren't living up to their potential. I argue that realism trumps optimism. It's perfectly normal in a realist farming to see something difficult, acknowledge the high risk and failure potential, and still pursue something with intent to succeed. I've personally grown tired of over optimism everywhere because it c…

Realism doesn't work in business. Business success requires 10 people to try for 1 person to succeed. If those 10 people were realists, they wouldn't try.

Depends how much that one person wins and how much the others lose.

Re: Show HN: San Francisco Compute – 512 H100s at <$2/hr for research and startups

#122
post #74
post #70

> Rather than each of K startups individually buying clusters of N gpus, together we buy a cluster with NK gpus... Then we set up a job scheduler to allocate compute In theory, this sounds almost identical to the business model behind AWS, Azure, and other cloud providers. "Instead of everyone buying a fixed amount of hardware for individual use, we'll buy a massive pool of hardware that people can time-share." Outsi…

AWS and Azure would slit their own throats before they created a way for their customers to pool instances to save money. They want to do that themselves, and keep the customer relationship and the profits, instead of giving them to a middleman or the customer.

AWS and Azure both charge by the hour anyway so it wouldn't but if you wanted you could use Reserved instances and just have their accounts in the same organisation.

A large part of the profit comes from the upfront risk of buying machines. With this you are just absorbing that risk which may be better if the startup expects to last.

Re: Show HN: San Francisco Compute – 512 H100s at <$2/hr for research and startups

#123
post #99

Earlier quoted context omitted.

Actually, the TPU Research Cloud program is still going strong! We've expanded the compute pool significantly to include Cloud TPU v4 Pod slices, and larger projects still use hundreds of chips at a time. (TRC capacity has not been reclaimed for internal use.) Check out this list of recent TRC-supported publications: https://sites.research.google/trc/publications/ Demand for Cloud TPUs is definitely intense, so if yo…

Zak, I love you buddy, but you should have some of your researchers try to use the TRC program. They should pretend to be a nobody (like I was in 2019) and try to do any research with the resources they’re granted. I guarantee you those researchers will all tell you “we can’t start any training runs anymore because the TPUs die after 45 minutes.” This may feel like an anime betrayal, since you basically launched my c…

A few quick comments:

> But it’s important for hobbyists and tinkerers to be able to participate in the AI ecosystem

Totally agree! This was a big part of my original motivation for creating the TPU Research Cloud program. People sometimes assume that e.g. an academic affiliation is required to participate, but that isn't true; we want the program to be as open as possible. We should find a better way to highlight the work of TRC tinkerers - for now, the GitHub and Hugging Face search buttons near the top of https://sites.research.google/trc/publications/ provide some raw pointers.

I'm sorry to hear that you've personally had a hard time getting TPU v3 capacity in europe-west4-a. In general, TRC TPU availability varies by region and by hardware generation, and we've experimented with different ways of prioritizing projects. It's possible that something was misconfigured on our end if your TPU lifetimes were so short. Could you email Jonathan the name of the project(s) you were using and any other data you still have handy so we can figure out what was going wrong?

Also, thanks for the kind words for Jonathan and the rest of the TRC team. They haven't lost any power or control, and they are allocating a lot more Cloud TPU capacity than ever. However, now that everyone wants to train LLMs, diffusion models, and other exciting new things, demand for TPU compute is way up, so juggling all of the inbound TRC requests is definitely more challenging than it used to be.

Re: Show HN: San Francisco Compute – 512 H100s at <$2/hr for research and startups

#124
post #67
post #30

Earlier quoted context omitted.

Optimism is (almost) always required in order to accomplish anything of significance. Those who lose it, aren't living up to their potential. I'm not encouraging the false belief that everything you do will work out. Instead I'm encouraging the realization that the greatest accomplishments almost always feel like long shots, and require significant amounts of optimism. Fear and pessimism, while helpful in appropriate…

>Optimism is (almost) always required in order to accomplish anything of significance. Those who lose it, aren't living up to their potential. I argue that realism trumps optimism. It's perfectly normal in a realist farming to see something difficult, acknowledge the high risk and failure potential, and still pursue something with intent to succeed. I've personally grown tired of over optimism everywhere because it c…

While I've said something like this comment scores of times in my life, and it's definitely a necessary corrective for a lot of optimists who don't think too hard about how they think, I don't think it's a useful place to stop. It's not hard to get unanimous agreement with "be a realist!" because it's framed so the alternative is irrationality/delusion. But even among people who agree that the goal should be to reason under uncertainty and assess risks clearly, there will be a spectrum of risk tolerance, and I don't think it's the worst thing ever to describe that as "optimism" vs. "pessimism"! (I fully acknowledge this isn't the dominant usage, but I think some spaces lean this way)

In this context, I tend to read the parent claim as something like, "great success requires willingness to sometimes take worse-than-even odds or pursue modestly-negative-EV opportunities". I'm not sure I agree with the strongest version of that, but I think it's likely that the space of risky paths to great achievement is richer than that of cautious ones.

Re: Show HN: San Francisco Compute – 512 H100s at <$2/hr for research and startups

#125
post #120

Earlier quoted context omitted.

Zak, I love you buddy, but you should have some of your researchers try to use the TRC program. They should pretend to be a nobody (like I was in 2019) and try to do any research with the resources they’re granted. I guarantee you those researchers will all tell you “we can’t start any training runs anymore because the TPUs die after 45 minutes.” This may feel like an anime betrayal, since you basically launched my c…

> You don’t need to take my word for it. Here’s some unfiltered DMs on the subject: https://imgur.com/a/6vqvzXs > Notice how their optimism dries up, and not because I was telling them how bad TRC has become. It’s because their TPUs kept dying. Unless I'm misreading this they sound pretty happy and you sound pessimistic? Their last substantial comment was "I'm sure Zak could hook you up with something better"?

TRC is supposed to be the “something better”. This insider TPU stuff is for the birds. If TRC can only offer 4 hours with no preemptions, that’s fine, but they need to be up front about that. Saying that TPUs preempt every 24 hours and then killing them off after 45 minutes is… not very productive.

As for their comments, the third screenshot is the key; they’re agreeing that the situation is bad. They’re a friend, and they’re a little indirect with the way they phrase things. (If you’ve ever had a friend who really doesn’t want to be wrong, you know what I mean; they kind of say things in a circular way in order to agree without agreeing. After awhile it’s pretty cute and endearing though.)

I was particularly pessimistic in those DMs because it came a couple months after I thought I’d give TRC one last try, back in January, which was roughly a year after I’d started my “ok, I’m losing hope, but I’ll wait and see” journey. In the meantime I kept cheerleading TRC and driving people to their signup page. But after the TPUs all died in less than two hours yet again, that was that.

I have a really high tolerance for faulty equipment. This is free compute; me complaining is just ungrateful. But I saw what things were like in 2019. “Different” would be the understatement of the century. If my baby wasn’t being incubated in the NICU today, I’d show the charts where our usage went from thousands of cores down to almost zero, and not for lack of trying.

It also would’ve been fine to say “sorry, this is unsustainable, the new limits are one tpu per person per project” and then give me a rock solid tpu. We had those in 2021. One of our TPUv3s stayed online for so long that I started to host my blog on it just to show people that TPUs were good for more than AI; the uptime was measured in months. Then poof, now you can barely fire one up.

Re: Show HN: San Francisco Compute – 512 H100s at <$2/hr for research and startups

#126

Earlier quoted context omitted.

Power seems like a very small amount of cost of compute when it comes to GPU’s.

FWIW I tired to look up some numbers, i found California "industrial" electricity at $0.18/Kwh https://www.eia.gov/electricity/monthly/epm_table_grapher.ph... and H100s using 300-700w https://www.nvidia.com/en-us/data-center/h100/ which implies a worst case marginal cost of .18*.7 = $.126 / gpu / hour. Looks like Montana is cheapest at ~$.05 / kwh which would bring that down to $.035. So there may be about a $0.09 Ca…

Retail residential power in the city of Santa Clara is $0.15/KwH, I'm sure commercial could be less. Especially if you throw some solar panels on the roof.

The most expensive part would be the land, but honestly there is some pretty cheap land outside the cities.

Re: Show HN: San Francisco Compute – 512 H100s at <$2/hr for research and startups

#127
post #40

Earlier quoted context omitted.

Oh no, definitely not. We just got a loan. Neither Alex or I are currently VCs, and this has no affiliation with any venture fund. We want to be a customer of the sf compute group too!

How’d you get this loan? Is it from a benevolent individual who just wants to make something happen? If not and you got the loan from a bank, super curious how you were able to get the bank to trust that renting out the GPUs would cover the loan or if some other reasoning convinced them. Assuming you aren’t trying to turn this into a big business, that knowledge might help a lot of other players run similar programs…

I’m fairly certain this loan is either a private individual or a HELOC or something. No way is a bank just going to loan out a bunch of money to some startup like this.

Re: Show HN: San Francisco Compute – 512 H100s at <$2/hr for research and startups

#128
post #67

Earlier quoted context omitted.

>Optimism is (almost) always required in order to accomplish anything of significance. Those who lose it, aren't living up to their potential. I argue that realism trumps optimism. It's perfectly normal in a realist farming to see something difficult, acknowledge the high risk and failure potential, and still pursue something with intent to succeed. I've personally grown tired of over optimism everywhere because it c…

If everyone were a realist, we wouldn't have half the advances we do. Because what can be "real" is proven wrong through innovation, after all isn't that disruption? :) Sam Altman talks about this quite frequently, that it's not intelligence or luck necessary for an enduring innovation. It is persistence in the face of inevitability, and a high tolerance for being proven wrong and still persisting

Everybody knows socialism is impossible though. Can't work, not worth trying, don't even think about it.

Re: Show HN: San Francisco Compute – 512 H100s at <$2/hr for research and startups

#129
post #123

Earlier quoted context omitted.

Zak, I love you buddy, but you should have some of your researchers try to use the TRC program. They should pretend to be a nobody (like I was in 2019) and try to do any research with the resources they’re granted. I guarantee you those researchers will all tell you “we can’t start any training runs anymore because the TPUs die after 45 minutes.” This may feel like an anime betrayal, since you basically launched my c…

A few quick comments: > But it’s important for hobbyists and tinkerers to be able to participate in the AI ecosystem Totally agree! This was a big part of my original motivation for creating the TPU Research Cloud program. People sometimes assume that e.g. an academic affiliation is required to participate, but that isn't true; we want the program to be as open as possible. We should find a better way to highlight th…

It’s not euw4a. It’s everywhere. The allocation algorithm across the board kills off TPUs after no more than a couple hours. usc1f, usc1a, usc1c, euw4a; it makes no difference.

It would be funny if someone set gpt-2-15b-poetry (our project) in some special way to prevent us from making TPUs that ever last more than a few hours, but from what I’ve heard from other people, this isn’t the case. That’s what I mean about the left hand doesn’t know what’s going on with the right hand. It’s not a misconfiguration. Again, pretend to be some random person who just wants to apply for TPU access, fill out your form, then try to do research with the TPUs that are available to you. You’ll have a rough time, but it’ll also cure this misconception that it’s a special case or was just me.

Again, no need to take my word for it; here’s an organic comment from someone who was rolling their eyes whenever I was cheerleading TRC, because their experience was so bad: https://news.ycombinator.com/item?id=36936782

I think that the experience is probably great for researchers who get special approval. And that’s fine, if that’s how the program is designed to be. But at least tell people that they shouldn’t expect more than an hour or two of TPU time.

Re: Show HN: San Francisco Compute – 512 H100s at <$2/hr for research and startups

#130
post #70

> Rather than each of K startups individually buying clusters of N gpus, together we buy a cluster with NK gpus... Then we set up a job scheduler to allocate compute In theory, this sounds almost identical to the business model behind AWS, Azure, and other cloud providers. "Instead of everyone buying a fixed amount of hardware for individual use, we'll buy a massive pool of hardware that people can time-share." Outsi…

Couple things, mostly pricing and availability:

1) Margins. Public cloud investors expect a certain margin profile. They can’t compete with Lambda/Fluidstack’s margins.

2) To an extent also big clouds have worse networking for LLM training. I believe only Azure has infiniband. Oracle is 3200 Gbps but not infiniband, same for AWS I believe. GCP not sure but their A100 networking speeds were only 100 Gbps I believe rather than 1600. Whereas lambda, fluidstack and coreweave all have ib.

3) Availability. Nvidia isn’t giving big clouds the allocation they want.

Post reply on HN