The Dojo is open.
I thought Dojo was custom chips.
Tesla turns on 10k-node Nvidia H100 Cluster
21–30 of 135 posts
Re: Tesla turns on 10k-node Nvidia H100 Cluster
#22I was curious why this statement lead with fp64 flops (instead of fp32, perhaps), but I looked up the H100 specs, and NV’s marketing page does the same thing. They’re obviously talking about the H100 SXM here, which has the same peak theoretical fp64 throughput as fp32. The cluster perf is estimated by multiplying the GPU perf by 10k.
Also, obviously, int8 tensor ops aren’t ‘FLOPS’. I think Nvidia calls them “TOPS” (tensor ops). There is a separate metric for ‘tensor flops’ or TF32.
Re: Tesla turns on 10k-node Nvidia H100 Cluster
#23I understand that the H100 is NVidia's leading edge chip, but can someone let me know if 10K is considered to be a big cluster? I've never worked inside one of the leading edge AI companies like OpenAI, Google, Microsoft or Meta. Is this comparable to what they would work with? My first guess is that it seems much smaller. And if you are running many parallel training jobs then you are getting about 1,000 chips at mo…
Re: Tesla turns on 10k-node Nvidia H100 Cluster
#24Nvidia is powering a mega Tesla supercomputer powered by 10,000 H100 GPUs
Re: Tesla turns on 10k-node Nvidia H100 Cluster
#25What happened to their custom hardware training stack Dojo? They had some interesting ideas there. The last I heard, they had one of those tiles "working" in the lab. Pretty far from a production setup. I can imagine they either underestimated the software effort needed to squeeze as much performance as possible out of those things, or they underestimated the pace at which Nvidia scales FLOPS/$, or both.
Re: Tesla turns on 10k-node Nvidia H100 Cluster
#26Earlier quoted context omitted.
It’s bottleneck on Nvidia side. They are producing less than Tesla consume. Tesla compute power will outclass many cloud provider combined in just three or four years with their own custom chip.
Got a source for that?
https://twitter.com/SawyerMerritt/status/1696011140508045660
Re: Tesla turns on 10k-node Nvidia H100 Cluster
#27What happened to their custom hardware training stack Dojo? They had some interesting ideas there. The last I heard, they had one of those tiles "working" in the lab. Pretty far from a production setup. I can imagine they either underestimated the software effort needed to squeeze as much performance as possible out of those things, or they underestimated the pace at which Nvidia scales FLOPS/$, or both.
Vaporware, just like much of what Musk talks about.
What would you say grants you the standing to opine here?
Re: Tesla turns on 10k-node Nvidia H100 Cluster
#28Earlier quoted context omitted.
It’s bottleneck on Nvidia side. They are producing less than Tesla consume. Tesla compute power will outclass many cloud provider combined in just three or four years with their own custom chip.
Got a source for that?
Re: Tesla turns on 10k-node Nvidia H100 Cluster
#29I understand that the H100 is NVidia's leading edge chip, but can someone let me know if 10K is considered to be a big cluster? I've never worked inside one of the leading edge AI companies like OpenAI, Google, Microsoft or Meta. Is this comparable to what they would work with? My first guess is that it seems much smaller. And if you are running many parallel training jobs then you are getting about 1,000 chips at mo…
Re: Tesla turns on 10k-node Nvidia H100 Cluster
#30Weird to think that his next company's compute platform is this.