Tesla turns on 10k-node Nvidia H100 Cluster
61–70 of 135 posts
Re: Tesla turns on 10k-node Nvidia H100 Cluster
#62Earlier quoted context omitted.
That Telsa owners can use their cars to make money while they are working as robo taxis -let's just say he vastly underestimates effort it takes to make progress - FSD is not there yet.
Vaporware assumes it will never happen. Is that the case you think or is it that he was vastly over optimistic? Very likely the latter.
Sucks to be if you paid the early bird fee for it.
Re: Tesla turns on 10k-node Nvidia H100 Cluster
#63n00b questions from someone just beginning to get interested in HPC I see mention of using this supercomputer for training models. Is that the only purpose? What other types of things do orgs usually do with these supercomputers? Are there any good boots-on-the-ground technical blogs that provide interesting detail on day-to-day experiences with these things?
In other words, they're used when you want to share some kind of state across all of the computers, without the potential overhead of communicating to some other system like a database.
Physics simulations and like, molecular modeling come to mind as common examples.
In the case of ML training, model parameters and broadcasting the deltas that get calculated during training are that shared state.
Re: Tesla turns on 10k-node Nvidia H100 Cluster
#64Earlier quoted context omitted.
That Telsa owners can use their cars to make money while they are working as robo taxis -let's just say he vastly underestimates effort it takes to make progress - FSD is not there yet.
Vaporware assumes it will never happen. Is that the case you think or is it that he was vastly over optimistic? Very likely the latter.
Re: Tesla turns on 10k-node Nvidia H100 Cluster
#65I understand that the H100 is NVidia's leading edge chip, but can someone let me know if 10K is considered to be a big cluster? I've never worked inside one of the leading edge AI companies like OpenAI, Google, Microsoft or Meta. Is this comparable to what they would work with? My first guess is that it seems much smaller. And if you are running many parallel training jobs then you are getting about 1,000 chips at mo…
10k H100 chips is considered a very large cluster. The third fastest supercomputer in the world is Microsoft’s eagle with 14k H100s https://www.top500.org/lists/top500/2023/11/
Re: Tesla turns on 10k-node Nvidia H100 Cluster
#66Earlier quoted context omitted.
Vaporware assumes it will never happen. Is that the case you think or is it that he was vastly over optimistic? Very likely the latter.
They will get there at some acceptable point but not with the tech in current Tesla's - the current compute module will need to be replaced - think they showed off HW 4 in lieu of HW 3. Sucks to be if you paid the early bird fee for it.
Re: Tesla turns on 10k-node Nvidia H100 Cluster
#67Earlier quoted context omitted.
The most powerful listed supercomputer has 37,888 Radeon GPUs, so in the same order of magnitude.
Interesting choice of words... I take you work for OpenAI? :) How large is their/'your' cluster? Probably the biggest in the world by now..
Re: Tesla turns on 10k-node Nvidia H100 Cluster
#68Earlier quoted context omitted.
Vaporware, just like much of what Musk talks about.
Reusable rockets, electric cars, solar panels... What would you say grants you the standing to opine here?
Re: Tesla turns on 10k-node Nvidia H100 Cluster
#69Earlier quoted context omitted.
Vaporware, just like much of what Musk talks about.
What that he has talked about been vaporware?