The Dojo is open.
Tesla turns on 10k-node Nvidia H100 Cluster
11–20 of 135 posts
Re: Tesla turns on 10k-node Nvidia H100 Cluster
#12I can imagine they either underestimated the software effort needed to squeeze as much performance as possible out of those things, or they underestimated the pace at which Nvidia scales FLOPS/$, or both.
Re: Tesla turns on 10k-node Nvidia H100 Cluster
#13Only 10K?
Re: Tesla turns on 10k-node Nvidia H100 Cluster
#14Only 10K?
Re: Tesla turns on 10k-node Nvidia H100 Cluster
#15I've never worked inside one of the leading edge AI companies like OpenAI, Google, Microsoft or Meta.
Is this comparable to what they would work with?
My first guess is that it seems much smaller. And if you are running many parallel training jobs then you are getting about 1,000 chips at most to work with.
Or is this about what the leading competitors are working with?
Azure, for one, seems to have orders of magnitude more chips at their disposal.
Re: Tesla turns on 10k-node Nvidia H100 Cluster
#16What happened to their custom hardware training stack Dojo? They had some interesting ideas there. The last I heard, they had one of those tiles "working" in the lab. Pretty far from a production setup. I can imagine they either underestimated the software effort needed to squeeze as much performance as possible out of those things, or they underestimated the pace at which Nvidia scales FLOPS/$, or both.
Re: Tesla turns on 10k-node Nvidia H100 Cluster
#17Original tweet: https://twitter.com/SawyerMerritt/status/1696011140508045660 Previus article: https://www.tomshardware.com/news/teslas-dollar300-million-a... This is second-hand blogspam.
Re: Tesla turns on 10k-node Nvidia H100 Cluster
#18Only 10K?
It’s bottleneck on Nvidia side. They are producing less than Tesla consume. Tesla compute power will outclass many cloud provider combined in just three or four years with their own custom chip.
That seems like a bold claim. Google, Microsoft and Meta make so much more money than Telsa that if making AI chips was so easy, then they could clearly out design and build Tesla without thinking too hard about it.
What makes you think that Telsa, a company with far less AI workers and knowledge, and far less money than the above companies can out design and out build them?
Re: Tesla turns on 10k-node Nvidia H100 Cluster
#19> The firm also built a compute cluster fitted with 5,760 Nvidia A100 GPUs in June 2012 Wow, that's some really early hardware access. /s