Live data from Hacker News

Tesla turns on 10k-node Nvidia H100 Cluster

techradar.com

121–130 of 135 posts

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#121

Earlier quoted context omitted.

I previously ran 150,000 AMD GPUs. 10k doesn't seem that large. =) That said, these GPUs aren't just the GPUs. They are whole chassis. They are huge onboard storage arrays, TB's of RAM, 800G networking (and associated cables), racks, cooling, power distribution, backup power, etc... None of it is easy.

H100 based DGX/HGX doesn't use 800 Gbit (it doesn't have the PCI-e bw), it's using 400 per GPU.

I was talking about between nodes. We're planning on bonding 2x400G NICs to get that 800G between nodes.

That said, latest 4th gen nvlink is 900G...

https://www.nvidia.com/en-us/data-center/nvlink/

But unless you're sleeping with Jensen, you're not going to see it for 52 weeks+ after you order it.

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#122
post #60

Earlier quoted context omitted.

In the old days, depending on architecture, fp64 performance could be atrocious even when fp32 was decent, so bragging about fp64 performance has an authenticity to it. Not all scientific computing requires 64 bits, but knowing that you can drop to high precision when necessary without penalty is nice. Also, back in the day, integer ops were just called 'ops', grumble grumble. But yeah FLOPS specifically refers to fl…

Still true that fp64 throughput is lower for consumer GPUs - both NV and AMD. That’s kinda why I was curious about leading with that metric - outside of HPC and scientific applications, a lot of people don’t really need fp64, and the machine might normally have a much higher fp32 throughput. > knowing you can drop to high precision when necessary without penalty is nice. I guess I maybe don’t know why you’d ever have…

You're right -- I have no idea why fp64 wouldn't be half the speed of fp32, and traditionally it is. I was simply taking them at their word. Maybe they're exaggerating or maybe they did what you suggest and hamstrung fp32.

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#124

I understand that the H100 is NVidia's leading edge chip, but can someone let me know if 10K is considered to be a big cluster? I've never worked inside one of the leading edge AI companies like OpenAI, Google, Microsoft or Meta. Is this comparable to what they would work with? My first guess is that it seems much smaller. And if you are running many parallel training jobs then you are getting about 1,000 chips at mo…

It's a small cluster the size of large cluster.

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#125
post #103

Earlier quoted context omitted.

Rocketry involves building a lot of prototypes and blowing up a lot of things when mistakes happen. Carmack did execute, they had a working rocket, and with time they could've solved the software issues, but they didn't have the funding to blow up dozens of prototypes. Carmack drove that project, he wrote code, he built rocket engines himself, he ran missions. Elon, notably, doesn't do any of that. He just had more m…

In general I think what is so often missed on HN, which is so ironic given this is in big part a board for founders or those who aspire to be (I thought), is the challenge and effort and success of building companies and teams that execute on larger than life projects. It is incredibly hard to recruit the best talent in the world, and assemble them and make it work. The leader must be incredibly competent and believa…

You left out a tier, and that's the embittered one with an axe to grind. It's the "nothing ever happens" crowd, and they want to destroy anyone who stands as a counterexample.

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#126

Earlier quoted context omitted.

Reusable rockets, electric cars, solar panels... What would you say grants you the standing to opine here?

https://elonmusk.today/

Most of those things have happened actually, but the website makes it seems that they didn't. It just lists everything Elon has said, but doesn't track if they happened or not. This is a completely pointless website.

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#127
post #80

Earlier quoted context omitted.

Out of interest, what did you use all that compute for?

ETH PoW. When ETH switched to PoS, we shut it all down. It sure was fun while it lasted, not many people on the planet have run that much compute. I did a lot of unique optimizations to autotune each individual GPU for performance by tweaking the software knobs on them. They are all snowflakes. Same model, different batches (heck, even same batch!), can produce wildly different performance results. Over the years, I…

Make a vid. Or a blog post, at least. Please :)

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#128
post #93

Earlier quoted context omitted.

Last I heard, the estimate was that NVIDIA would build 550k units in 2023, so 2% of all production — especially as at least six others (your four plus Apple and at least one intelligence agency) will be of similar size by themselves — is certainly non-negligible.

550k H100s? Who is buying these? They are hella expensive and China isn't allowed to have them.

Government agencies.

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#129

Earlier quoted context omitted.

Reusable rockets, electric cars, solar panels... What would you say grants you the standing to opine here?

All of his other false or misleading statements over the last 10 years.

When we've dealt with the oil companies, the chemical manufacturers dumping PFAS into our kids, and the industrial war machine, maybe then we can start complaining about the guy biting off more than he can chew trying to be constructive.

Until then, all of you sound vicious, bitter, and hypocritical.

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#130
post #127

Earlier quoted context omitted.

ETH PoW. When ETH switched to PoS, we shut it all down. It sure was fun while it lasted, not many people on the planet have run that much compute. I did a lot of unique optimizations to autotune each individual GPU for performance by tweaking the software knobs on them. They are all snowflakes. Same model, different batches (heck, even same batch!), can produce wildly different performance results. Over the years, I…

Make a vid. Or a blog post, at least. Please :)

Thanks, but not my style, sorry! I've been doing PoW mining since 2014 and have so many stories, I've forgotten half of them. I wouldn't even know where to start on trying to document any of it.
Post reply on HN