Live data from Hacker News

Tesla turns on 10k-node Nvidia H100 Cluster

techradar.com

111–120 of 135 posts

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#111
post #93

Earlier quoted context omitted.

Last I heard, the estimate was that NVIDIA would build 550k units in 2023, so 2% of all production — especially as at least six others (your four plus Apple and at least one intelligence agency) will be of similar size by themselves — is certainly non-negligible.

550k H100s? Who is buying these? They are hella expensive and China isn't allowed to have them.

The Big Cloud

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#112

Earlier quoted context omitted.

ETH PoW. When ETH switched to PoS, we shut it all down. It sure was fun while it lasted, not many people on the planet have run that much compute. I did a lot of unique optimizations to autotune each individual GPU for performance by tweaking the software knobs on them. They are all snowflakes. Same model, different batches (heck, even same batch!), can produce wildly different performance results. Over the years, I…

Did you manage to recoup the investment?

Of course I can't say anything about that other than I did the job I was hired to do, and I performed far above anyone's wildest expectations.

Nobody else on the planet was able to automate the tuning like I did, which had a direct influence on ROI. I know this because it required a very specific change to the AMD drivers to enable that functionality to happen.

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#113

I understand that the H100 is NVidia's leading edge chip, but can someone let me know if 10K is considered to be a big cluster? I've never worked inside one of the leading edge AI companies like OpenAI, Google, Microsoft or Meta. Is this comparable to what they would work with? My first guess is that it seems much smaller. And if you are running many parallel training jobs then you are getting about 1,000 chips at mo…

I previously ran 150,000 AMD GPUs. 10k doesn't seem that large. =) That said, these GPUs aren't just the GPUs. They are whole chassis. They are huge onboard storage arrays, TB's of RAM, 800G networking (and associated cables), racks, cooling, power distribution, backup power, etc... None of it is easy.

H100 based DGX/HGX doesn't use 800 Gbit (it doesn't have the PCI-e bw), it's using 400 per GPU.

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#116

Earlier quoted context omitted.

Harder for everyone, including his staff, who were asking him to move to Windows…

He should have hired staff that is competent with the tech stack used at his company. Unforced rewrites are usually always a bad idea.

Oh that simple, huh? Too bad you weren’t around in the late 90s to explain this to him, and to help him find the extremely rare group of folks familiar with Linux…

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#117

Earlier quoted context omitted.

Harder for who? Elon certainly didn't have the technical chops to work with it.

Harder for everyone, including his staff, who were asking him to move to Windows…

Actually he was the other way around. He strongly pushing windows and the CTO and engineers from Confonity strongly wanted Linux

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#118

Earlier quoted context omitted.

Harder for everyone, including his staff, who were asking him to move to Windows…

Actually he was the other way around. He strongly pushing windows and the CTO and engineers from Confonity strongly wanted Linux

“Wanting to rewrite for windows” is what you said; was that not accurate?

The key thing here however is that Elon didn’t want whatever he asked for in a vacuum, despite what the book says. Surely this was his engineer’s preference.

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#119
post #93

Earlier quoted context omitted.

Last I heard, the estimate was that NVIDIA would build 550k units in 2023, so 2% of all production — especially as at least six others (your four plus Apple and at least one intelligence agency) will be of similar size by themselves — is certainly non-negligible.

550k H100s? Who is buying these? They are hella expensive and China isn't allowed to have them.

Other than the ~12% I just estimated, lots of large-but-not-famous places will be buying ~1k, and small places will be buying tens to hundreds, and quite a lot of AI bubble money will be invested in startups that claim they only need one.

Probably some scientific modelling that can be done on these, so I bet some universities and private labs will be buying them. NASA, SpaceX, RocketLab, Helion, etc.

There's also probably a lot of AAA game studios and art studios for movies etc. who are each buying dozens of these graphics processing units for… graphics :P

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#120
post #29

Earlier quoted context omitted.

The most powerful listed supercomputer has 37,888 Radeon GPUs, so in the same order of magnitude.

Interesting choice of words... I take you work for OpenAI? :) How large is their/'your' cluster? Probably the biggest in the world by now..

Unfortunately no, but there are almost certainly clusters in the hands of private companies and government organizations that would prefer not to advertise their capabilities.
Post reply on HN