Live data from Hacker News

Tesla turns on 10k-node Nvidia H100 Cluster

techradar.com

21–30 of 135 posts

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#21

The Dojo is open.

I thought Dojo was custom chips.

You are correct; it is, and flippant HN comments that are additionally incorrect are starting to become a thing. See the original tweet: https://twitter.com/SawyerMerritt/status/1696011140508045660

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#22
> This AI cluster, worth more than $300 million, will offer a peak performance of 340 FP64 PFLOPS for technical computing and 39.58 INT8 ExaFLOPS for AI applications, according to Tom’s Hardware.

I was curious why this statement lead with fp64 flops (instead of fp32, perhaps), but I looked up the H100 specs, and NV’s marketing page does the same thing. They’re obviously talking about the H100 SXM here, which has the same peak theoretical fp64 throughput as fp32. The cluster perf is estimated by multiplying the GPU perf by 10k.

Also, obviously, int8 tensor ops aren’t ‘FLOPS’. I think Nvidia calls them “TOPS” (tensor ops). There is a separate metric for ‘tensor flops’ or TF32.

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#23

I understand that the H100 is NVidia's leading edge chip, but can someone let me know if 10K is considered to be a big cluster? I've never worked inside one of the leading edge AI companies like OpenAI, Google, Microsoft or Meta. Is this comparable to what they would work with? My first guess is that it seems much smaller. And if you are running many parallel training jobs then you are getting about 1,000 chips at mo…

[dead]

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#25
post #12

What happened to their custom hardware training stack Dojo? They had some interesting ideas there. The last I heard, they had one of those tiles "working" in the lab. Pretty far from a production setup. I can imagine they either underestimated the software effort needed to squeeze as much performance as possible out of those things, or they underestimated the pace at which Nvidia scales FLOPS/$, or both.

Probably they want to have all and any compute they can have. This doesn't exclude Dojo nor the previous generation nvidia chips they already got.

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#26

Earlier quoted context omitted.

It’s bottleneck on Nvidia side. They are producing less than Tesla consume. Tesla compute power will outclass many cloud provider combined in just three or four years with their own custom chip.

Got a source for that?

The original tweet makes the claim, but the tweet seems prone to hyperbole as well.

https://twitter.com/SawyerMerritt/status/1696011140508045660

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#27
post #12

What happened to their custom hardware training stack Dojo? They had some interesting ideas there. The last I heard, they had one of those tiles "working" in the lab. Pretty far from a production setup. I can imagine they either underestimated the software effort needed to squeeze as much performance as possible out of those things, or they underestimated the pace at which Nvidia scales FLOPS/$, or both.

Vaporware, just like much of what Musk talks about.

Reusable rockets, electric cars, solar panels...

What would you say grants you the standing to opine here?

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#28

Earlier quoted context omitted.

It’s bottleneck on Nvidia side. They are producing less than Tesla consume. Tesla compute power will outclass many cloud provider combined in just three or four years with their own custom chip.

Got a source for that?

The original tweet quotes Elon Musk saying "Frankly...if they (NVIDIA) could deliver us enough GPUs, we might not need Dojo"

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#29

I understand that the H100 is NVidia's leading edge chip, but can someone let me know if 10K is considered to be a big cluster? I've never worked inside one of the leading edge AI companies like OpenAI, Google, Microsoft or Meta. Is this comparable to what they would work with? My first guess is that it seems much smaller. And if you are running many parallel training jobs then you are getting about 1,000 chips at mo…

The most powerful listed supercomputer has 37,888 Radeon GPUs, so in the same order of magnitude.
Post reply on HN