I understand that the H100 is NVidia's leading edge chip, but can someone let me know if 10K is considered to be a big cluster? I've never worked inside one of the leading edge AI companies like OpenAI, Google, Microsoft or Meta. Is this comparable to what they would work with? My first guess is that it seems much smaller. And if you are running many parallel training jobs then you are getting about 1,000 chips at mo…
Tesla turns on 10k-node Nvidia H100 Cluster
31–40 of 135 posts
Re: Tesla turns on 10k-node Nvidia H100 Cluster
#32It's funny - I'm listening to "The Founders" audiobook and right now they're telling the story of Elon Musk at PayPal wanting to rewrite for Windows server because Linux was too hard. Weird to think that his next company's compute platform is this.
Re: Tesla turns on 10k-node Nvidia H100 Cluster
#33Re: Tesla turns on 10k-node Nvidia H100 Cluster
#34> This AI cluster, worth more than $300 million, will offer a peak performance of 340 FP64 PFLOPS for technical computing and 39.58 INT8 ExaFLOPS for AI applications, according to Tom’s Hardware. I was curious why this statement lead with fp64 flops (instead of fp32, perhaps), but I looked up the H100 specs, and NV’s marketing page does the same thing. They’re obviously talking about the H100 SXM here, which has the…
Re: Tesla turns on 10k-node Nvidia H100 Cluster
#35> The firm also built a compute cluster fitted with 5,760 Nvidia A100 GPUs in June 2012 Wow, that's some really early hardware access. /s
Re: Tesla turns on 10k-node Nvidia H100 Cluster
#36> The firm also built a compute cluster fitted with 5,760 Nvidia A100 GPUs in June 2012 Wow, that's some really early hardware access. /s
I only had about 3 NVIDIA H100 in 1980
Re: Tesla turns on 10k-node Nvidia H100 Cluster
#37What happened to their custom hardware training stack Dojo? They had some interesting ideas there. The last I heard, they had one of those tiles "working" in the lab. Pretty far from a production setup. I can imagine they either underestimated the software effort needed to squeeze as much performance as possible out of those things, or they underestimated the pace at which Nvidia scales FLOPS/$, or both.
Vaporware, just like much of what Musk talks about.
Re: Tesla turns on 10k-node Nvidia H100 Cluster
#38Re: Tesla turns on 10k-node Nvidia H100 Cluster
#39Re: Tesla turns on 10k-node Nvidia H100 Cluster
#40Earlier quoted context omitted.
I thought Dojo was custom chips.
You are correct; it is, and flippant HN comments that are additionally incorrect are starting to become a thing. See the original tweet: https://twitter.com/SawyerMerritt/status/1696011140508045660
https://en.wikipedia.org/wiki/Tesla_Dojo
From the History section (although Technical Architecure is also worthy of consuming in its entirety):
> In August 2023, Tesla powered on Dojo for production use as well as a new training cluster configured with 10,000 Nvidia H100 GPUs.
I’ll take the L wrt being flippant if we’re using words very specifically in this context, that’s fair. It’s great to see Tesla expand its training resources is my sentiment, regardless of how their aggregate ML compute is segmented.