Live data from Hacker News

Tesla turns on 10k-node Nvidia H100 Cluster

techradar.com

11–20 of 135 posts

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#12
What happened to their custom hardware training stack Dojo? They had some interesting ideas there. The last I heard, they had one of those tiles "working" in the lab. Pretty far from a production setup.

I can imagine they either underestimated the software effort needed to squeeze as much performance as possible out of those things, or they underestimated the pace at which Nvidia scales FLOPS/$, or both.

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#15
I understand that the H100 is NVidia's leading edge chip, but can someone let me know if 10K is considered to be a big cluster?

I've never worked inside one of the leading edge AI companies like OpenAI, Google, Microsoft or Meta.

Is this comparable to what they would work with?

My first guess is that it seems much smaller. And if you are running many parallel training jobs then you are getting about 1,000 chips at most to work with.

Or is this about what the leading competitors are working with?

Azure, for one, seems to have orders of magnitude more chips at their disposal.

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#16
post #12

What happened to their custom hardware training stack Dojo? They had some interesting ideas there. The last I heard, they had one of those tiles "working" in the lab. Pretty far from a production setup. I can imagine they either underestimated the software effort needed to squeeze as much performance as possible out of those things, or they underestimated the pace at which Nvidia scales FLOPS/$, or both.

Vaporware, just like much of what Musk talks about.

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#17
post #5

Original tweet: https://twitter.com/SawyerMerritt/status/1696011140508045660 Previus article: https://www.tomshardware.com/news/teslas-dollar300-million-a... This is second-hand blogspam.

And the original tweet is very much kool-aid heavy, with "20x performance", "30x performance" claims about the switch from one card to the next.

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#18
post #6

Only 10K?

It’s bottleneck on Nvidia side. They are producing less than Tesla consume. Tesla compute power will outclass many cloud provider combined in just three or four years with their own custom chip.

> It’s bottleneck on Nvidia side. They are producing less than Tesla consume. Tesla compute power will outclass many cloud provider combined in just three or four years with their own custom chip.

That seems like a bold claim. Google, Microsoft and Meta make so much more money than Telsa that if making AI chips was so easy, then they could clearly out design and build Tesla without thinking too hard about it.

What makes you think that Telsa, a company with far less AI workers and knowledge, and far less money than the above companies can out design and out build them?

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#20
post #6

Only 10K?

It’s bottleneck on Nvidia side. They are producing less than Tesla consume. Tesla compute power will outclass many cloud provider combined in just three or four years with their own custom chip.

Got a source for that?
Post reply on HN