Live data from Hacker News

Tesla turns on 10k-node Nvidia H100 Cluster

techradar.com

31–40 of 135 posts

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#31

I understand that the H100 is NVidia's leading edge chip, but can someone let me know if 10K is considered to be a big cluster? I've never worked inside one of the leading edge AI companies like OpenAI, Google, Microsoft or Meta. Is this comparable to what they would work with? My first guess is that it seems much smaller. And if you are running many parallel training jobs then you are getting about 1,000 chips at mo…

10k H100 chips is considered a very large cluster. The third fastest supercomputer in the world is Microsoft’s eagle with 14k H100s https://www.top500.org/lists/top500/2023/11/

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#32

It's funny - I'm listening to "The Founders" audiobook and right now they're telling the story of Elon Musk at PayPal wanting to rewrite for Windows server because Linux was too hard. Weird to think that his next company's compute platform is this.

Linux was a lot harder back then.

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#33

Earlier quoted context omitted.

Vaporware, just like much of what Musk talks about.

Reusable rockets, electric cars, solar panels... What would you say grants you the standing to opine here?

All of his other false or misleading statements over the last 10 years.

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#34
post #22

> This AI cluster, worth more than $300 million, will offer a peak performance of 340 FP64 PFLOPS for technical computing and 39.58 INT8 ExaFLOPS for AI applications, according to Tom’s Hardware. I was curious why this statement lead with fp64 flops (instead of fp32, perhaps), but I looked up the H100 specs, and NV’s marketing page does the same thing. They’re obviously talking about the H100 SXM here, which has the…

Nit: INT8 is not a floating point operation and thus cannot be used in the term "ExaFLOPS"

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#36
post #19
post #7

> The firm also built a compute cluster fitted with 5,760 Nvidia A100 GPUs in June 2012 Wow, that's some really early hardware access. /s

I only had about 3 NVIDIA H100 in 1980

Someone needs to figure out at what point all the compute in the world became more powerful than a single H100.

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#37
post #12

What happened to their custom hardware training stack Dojo? They had some interesting ideas there. The last I heard, they had one of those tiles "working" in the lab. Pretty far from a production setup. I can imagine they either underestimated the software effort needed to squeeze as much performance as possible out of those things, or they underestimated the pace at which Nvidia scales FLOPS/$, or both.

Vaporware, just like much of what Musk talks about.

Actually they're in the middle of production at TSMC. They have 10,000 units on order, to be delivered "in the coming year".

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#38

Earlier quoted context omitted.

Vaporware, just like much of what Musk talks about.

Reusable rockets, electric cars, solar panels... What would you say grants you the standing to opine here?

I'm fairly certain all of those existed prior to Musk's suggestion of them.

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#40

Earlier quoted context omitted.

I thought Dojo was custom chips.

You are correct; it is, and flippant HN comments that are additionally incorrect are starting to become a thing. See the original tweet: https://twitter.com/SawyerMerritt/status/1696011140508045660

You’re being pedantic (rightfully so) and I’m being loose with words. While Dojo is the supercomputer Tesla built for vision training, I lumped anything contributing to their machine vision model training as Dojo. It’s called Dojo because that’s where the training takes place.

https://en.wikipedia.org/wiki/Tesla_Dojo

From the History section (although Technical Architecure is also worthy of consuming in its entirety):

> In August 2023, Tesla powered on Dojo for production use as well as a new training cluster configured with 10,000 Nvidia H100 GPUs.

I’ll take the L wrt being flippant if we’re using words very specifically in this context, that’s fair. It’s great to see Tesla expand its training resources is my sentiment, regardless of how their aggregate ML compute is segmented.

Post reply on HN