Earlier quoted context omitted.
>Optimism is (almost) always required in order to accomplish anything of significance. Those who lose it, aren't living up to their potential. I argue that realism trumps optimism. It's perfectly normal in a realist farming to see something difficult, acknowledge the high risk and failure potential, and still pursue something with intent to succeed. I've personally grown tired of over optimism everywhere because it c…
Realism doesn't work in business. Business success requires 10 people to try for 1 person to succeed. If those 10 people were realists, they wouldn't try.
Show HN: San Francisco Compute – 512 H100s at <$2/hr for research and startups
121–130 of 189 posts
Re: Show HN: San Francisco Compute – 512 H100s at <$2/hr for research and startups
#122> Rather than each of K startups individually buying clusters of N gpus, together we buy a cluster with NK gpus... Then we set up a job scheduler to allocate compute In theory, this sounds almost identical to the business model behind AWS, Azure, and other cloud providers. "Instead of everyone buying a fixed amount of hardware for individual use, we'll buy a massive pool of hardware that people can time-share." Outsi…
AWS and Azure would slit their own throats before they created a way for their customers to pool instances to save money. They want to do that themselves, and keep the customer relationship and the profits, instead of giving them to a middleman or the customer.
A large part of the profit comes from the upfront risk of buying machines. With this you are just absorbing that risk which may be better if the startup expects to last.
Re: Show HN: San Francisco Compute – 512 H100s at <$2/hr for research and startups
#123Earlier quoted context omitted.
Actually, the TPU Research Cloud program is still going strong! We've expanded the compute pool significantly to include Cloud TPU v4 Pod slices, and larger projects still use hundreds of chips at a time. (TRC capacity has not been reclaimed for internal use.) Check out this list of recent TRC-supported publications: https://sites.research.google/trc/publications/ Demand for Cloud TPUs is definitely intense, so if yo…
Zak, I love you buddy, but you should have some of your researchers try to use the TRC program. They should pretend to be a nobody (like I was in 2019) and try to do any research with the resources they’re granted. I guarantee you those researchers will all tell you “we can’t start any training runs anymore because the TPUs die after 45 minutes.” This may feel like an anime betrayal, since you basically launched my c…
> But it’s important for hobbyists and tinkerers to be able to participate in the AI ecosystem
Totally agree! This was a big part of my original motivation for creating the TPU Research Cloud program. People sometimes assume that e.g. an academic affiliation is required to participate, but that isn't true; we want the program to be as open as possible. We should find a better way to highlight the work of TRC tinkerers - for now, the GitHub and Hugging Face search buttons near the top of https://sites.research.google/trc/publications/ provide some raw pointers.
I'm sorry to hear that you've personally had a hard time getting TPU v3 capacity in europe-west4-a. In general, TRC TPU availability varies by region and by hardware generation, and we've experimented with different ways of prioritizing projects. It's possible that something was misconfigured on our end if your TPU lifetimes were so short. Could you email Jonathan the name of the project(s) you were using and any other data you still have handy so we can figure out what was going wrong?
Also, thanks for the kind words for Jonathan and the rest of the TRC team. They haven't lost any power or control, and they are allocating a lot more Cloud TPU capacity than ever. However, now that everyone wants to train LLMs, diffusion models, and other exciting new things, demand for TPU compute is way up, so juggling all of the inbound TRC requests is definitely more challenging than it used to be.
Re: Show HN: San Francisco Compute – 512 H100s at <$2/hr for research and startups
#124Earlier quoted context omitted.
Optimism is (almost) always required in order to accomplish anything of significance. Those who lose it, aren't living up to their potential. I'm not encouraging the false belief that everything you do will work out. Instead I'm encouraging the realization that the greatest accomplishments almost always feel like long shots, and require significant amounts of optimism. Fear and pessimism, while helpful in appropriate…
>Optimism is (almost) always required in order to accomplish anything of significance. Those who lose it, aren't living up to their potential. I argue that realism trumps optimism. It's perfectly normal in a realist farming to see something difficult, acknowledge the high risk and failure potential, and still pursue something with intent to succeed. I've personally grown tired of over optimism everywhere because it c…
In this context, I tend to read the parent claim as something like, "great success requires willingness to sometimes take worse-than-even odds or pursue modestly-negative-EV opportunities". I'm not sure I agree with the strongest version of that, but I think it's likely that the space of risky paths to great achievement is richer than that of cautious ones.
Re: Show HN: San Francisco Compute – 512 H100s at <$2/hr for research and startups
#125Earlier quoted context omitted.
Zak, I love you buddy, but you should have some of your researchers try to use the TRC program. They should pretend to be a nobody (like I was in 2019) and try to do any research with the resources they’re granted. I guarantee you those researchers will all tell you “we can’t start any training runs anymore because the TPUs die after 45 minutes.” This may feel like an anime betrayal, since you basically launched my c…
> You don’t need to take my word for it. Here’s some unfiltered DMs on the subject: https://imgur.com/a/6vqvzXs > Notice how their optimism dries up, and not because I was telling them how bad TRC has become. It’s because their TPUs kept dying. Unless I'm misreading this they sound pretty happy and you sound pessimistic? Their last substantial comment was "I'm sure Zak could hook you up with something better"?
As for their comments, the third screenshot is the key; they’re agreeing that the situation is bad. They’re a friend, and they’re a little indirect with the way they phrase things. (If you’ve ever had a friend who really doesn’t want to be wrong, you know what I mean; they kind of say things in a circular way in order to agree without agreeing. After awhile it’s pretty cute and endearing though.)
I was particularly pessimistic in those DMs because it came a couple months after I thought I’d give TRC one last try, back in January, which was roughly a year after I’d started my “ok, I’m losing hope, but I’ll wait and see” journey. In the meantime I kept cheerleading TRC and driving people to their signup page. But after the TPUs all died in less than two hours yet again, that was that.
I have a really high tolerance for faulty equipment. This is free compute; me complaining is just ungrateful. But I saw what things were like in 2019. “Different” would be the understatement of the century. If my baby wasn’t being incubated in the NICU today, I’d show the charts where our usage went from thousands of cores down to almost zero, and not for lack of trying.
It also would’ve been fine to say “sorry, this is unsustainable, the new limits are one tpu per person per project” and then give me a rock solid tpu. We had those in 2021. One of our TPUv3s stayed online for so long that I started to host my blog on it just to show people that TPUs were good for more than AI; the uptime was measured in months. Then poof, now you can barely fire one up.
Re: Show HN: San Francisco Compute – 512 H100s at <$2/hr for research and startups
#126Earlier quoted context omitted.
Power seems like a very small amount of cost of compute when it comes to GPU’s.
FWIW I tired to look up some numbers, i found California "industrial" electricity at $0.18/Kwh https://www.eia.gov/electricity/monthly/epm_table_grapher.ph... and H100s using 300-700w https://www.nvidia.com/en-us/data-center/h100/ which implies a worst case marginal cost of .18*.7 = $.126 / gpu / hour. Looks like Montana is cheapest at ~$.05 / kwh which would bring that down to $.035. So there may be about a $0.09 Ca…
The most expensive part would be the land, but honestly there is some pretty cheap land outside the cities.
Re: Show HN: San Francisco Compute – 512 H100s at <$2/hr for research and startups
#127Earlier quoted context omitted.
Oh no, definitely not. We just got a loan. Neither Alex or I are currently VCs, and this has no affiliation with any venture fund. We want to be a customer of the sf compute group too!
How’d you get this loan? Is it from a benevolent individual who just wants to make something happen? If not and you got the loan from a bank, super curious how you were able to get the bank to trust that renting out the GPUs would cover the loan or if some other reasoning convinced them. Assuming you aren’t trying to turn this into a big business, that knowledge might help a lot of other players run similar programs…
Re: Show HN: San Francisco Compute – 512 H100s at <$2/hr for research and startups
#128Earlier quoted context omitted.
>Optimism is (almost) always required in order to accomplish anything of significance. Those who lose it, aren't living up to their potential. I argue that realism trumps optimism. It's perfectly normal in a realist farming to see something difficult, acknowledge the high risk and failure potential, and still pursue something with intent to succeed. I've personally grown tired of over optimism everywhere because it c…
If everyone were a realist, we wouldn't have half the advances we do. Because what can be "real" is proven wrong through innovation, after all isn't that disruption? :) Sam Altman talks about this quite frequently, that it's not intelligence or luck necessary for an enduring innovation. It is persistence in the face of inevitability, and a high tolerance for being proven wrong and still persisting
Re: Show HN: San Francisco Compute – 512 H100s at <$2/hr for research and startups
#129Earlier quoted context omitted.
Zak, I love you buddy, but you should have some of your researchers try to use the TRC program. They should pretend to be a nobody (like I was in 2019) and try to do any research with the resources they’re granted. I guarantee you those researchers will all tell you “we can’t start any training runs anymore because the TPUs die after 45 minutes.” This may feel like an anime betrayal, since you basically launched my c…
A few quick comments: > But it’s important for hobbyists and tinkerers to be able to participate in the AI ecosystem Totally agree! This was a big part of my original motivation for creating the TPU Research Cloud program. People sometimes assume that e.g. an academic affiliation is required to participate, but that isn't true; we want the program to be as open as possible. We should find a better way to highlight th…
It would be funny if someone set gpt-2-15b-poetry (our project) in some special way to prevent us from making TPUs that ever last more than a few hours, but from what I’ve heard from other people, this isn’t the case. That’s what I mean about the left hand doesn’t know what’s going on with the right hand. It’s not a misconfiguration. Again, pretend to be some random person who just wants to apply for TPU access, fill out your form, then try to do research with the TPUs that are available to you. You’ll have a rough time, but it’ll also cure this misconception that it’s a special case or was just me.
Again, no need to take my word for it; here’s an organic comment from someone who was rolling their eyes whenever I was cheerleading TRC, because their experience was so bad: https://news.ycombinator.com/item?id=36936782
I think that the experience is probably great for researchers who get special approval. And that’s fine, if that’s how the program is designed to be. But at least tell people that they shouldn’t expect more than an hour or two of TPU time.
Re: Show HN: San Francisco Compute – 512 H100s at <$2/hr for research and startups
#130> Rather than each of K startups individually buying clusters of N gpus, together we buy a cluster with NK gpus... Then we set up a job scheduler to allocate compute In theory, this sounds almost identical to the business model behind AWS, Azure, and other cloud providers. "Instead of everyone buying a fixed amount of hardware for individual use, we'll buy a massive pool of hardware that people can time-share." Outsi…
1) Margins. Public cloud investors expect a certain margin profile. They can’t compete with Lambda/Fluidstack’s margins.
2) To an extent also big clouds have worse networking for LLM training. I believe only Azure has infiniband. Oracle is 3200 Gbps but not infiniband, same for AWS I believe. GCP not sure but their A100 networking speeds were only 100 Gbps I believe rather than 1600. Whereas lambda, fluidstack and coreweave all have ib.
3) Availability. Nvidia isn’t giving big clouds the allocation they want.