Live data from Hacker News

Tesla turns on 10k-node Nvidia H100 Cluster

techradar.com

81–90 of 135 posts

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#81
post #54
post #12

What happened to their custom hardware training stack Dojo? They had some interesting ideas there. The last I heard, they had one of those tiles "working" in the lab. Pretty far from a production setup. I can imagine they either underestimated the software effort needed to squeeze as much performance as possible out of those things, or they underestimated the pace at which Nvidia scales FLOPS/$, or both.

Dojo has always been a lie.

Your assertion is inaccurate.

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#82

Earlier quoted context omitted.

10k H100 chips is considered a very large cluster. The third fastest supercomputer in the world is Microsoft’s eagle with 14k H100s https://www.top500.org/lists/top500/2023/11/

Ah, gotcha, so the fact that its 10,000 chips for one dedicated cluster that makes it large, as opposed to Azure which has an order of magnitude more GPUS but rents many of those out.

I guess Azure's are spread out too. Latency higher to world wide datacentres.

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#83
post #41

Earlier quoted context omitted.

> It’s bottleneck on Nvidia side. They are producing less than Tesla consume. Tesla compute power will outclass many cloud provider combined in just three or four years with their own custom chip. That seems like a bold claim. Google, Microsoft and Meta make so much more money than Telsa that if making AI chips was so easy, then they could clearly out design and build Tesla without thinking too hard about it. What ma…

> What makes you think that Telsa, a company with far less AI workers and knowledge, an far less money than the above companies can out design and out build them? Presumably because Elon himself will be involved in the design, and Elon, as we all know, is one of the world's great thinkers. ;)

Elon is one of the worlds greatest talent poachers, and that is much better than being a great thinker.

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#84
post #49

Earlier quoted context omitted.

I'm fairly certain all of those existed prior to Musk's suggestion of them.

He delivered on them though right? Also reusable rockets didn’t exist?

John Carmack on a shoestring budget nearly got this working at Armadillo. If he had more money and time he would've had it working a half decade before SpaceX.

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#85
post #49

Earlier quoted context omitted.

He delivered on them though right? Also reusable rockets didn’t exist?

John Carmack on a shoestring budget nearly got this working at Armadillo. If he had more money and time he would've had it working a half decade before SpaceX.

Got it, so no then right.

I love Carmack, but having an idea for something is infinitely easier than actually successfully executing on it, especially as incredibly successfully as SpaceX.

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#86
post #80

Earlier quoted context omitted.

I previously ran 150,000 AMD GPUs. 10k doesn't seem that large. =) That said, these GPUs aren't just the GPUs. They are whole chassis. They are huge onboard storage arrays, TB's of RAM, 800G networking (and associated cables), racks, cooling, power distribution, backup power, etc... None of it is easy.

Out of interest, what did you use all that compute for?

Classified I imagine.

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#87
post #70

Newbie question, could this cluster easily calculate the largest prime number? I've found that the largest known prime number was found back in 2018, which is a while back considering how compute has evolved.

Finding the largest prime is more a contest of who's willing to commit the most ridiculous amount of compute to the goal than it is a mathematical obstacle.

The cost of finding the next prime is likely into the millions now.

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#88
post #55

Earlier quoted context omitted.

Vaporware assumes it will never happen. Is that the case you think or is it that he was vastly over optimistic? Very likely the latter.

FSD is currently in the quantum valley of product development: it is both vaporware and a shipping product.

The shipping product is FSD by name only. Actual automomy anywhere near the levels that has been promised to arrive "by the end of this year" for years will surely arrive by the end of this year.

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#89
post #80

Earlier quoted context omitted.

I previously ran 150,000 AMD GPUs. 10k doesn't seem that large. =) That said, these GPUs aren't just the GPUs. They are whole chassis. They are huge onboard storage arrays, TB's of RAM, 800G networking (and associated cables), racks, cooling, power distribution, backup power, etc... None of it is easy.

Out of interest, what did you use all that compute for?

ETH PoW. When ETH switched to PoS, we shut it all down. It sure was fun while it lasted, not many people on the planet have run that much compute.

I did a lot of unique optimizations to autotune each individual GPU for performance by tweaking the software knobs on them. They are all snowflakes. Same model, different batches (heck, even same batch!), can produce wildly different performance results.

Over the years, I did try to find some alternative workloads for it, but nothing could even pay for the power costs. The GPUs were very old models (rx470-rx580) and the rest of the hardware wasn't that advanced, like it is in AI, so none of it was transferred.

I'm in the process of building my own AI supercomputer now. Really looking forward to seeing how it turns out.

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#90
post #85

Earlier quoted context omitted.

John Carmack on a shoestring budget nearly got this working at Armadillo. If he had more money and time he would've had it working a half decade before SpaceX.

Got it, so no then right. I love Carmack, but having an idea for something is infinitely easier than actually successfully executing on it, especially as incredibly successfully as SpaceX.

Rocketry involves building a lot of prototypes and blowing up a lot of things when mistakes happen.

Carmack did execute, they had a working rocket, and with time they could've solved the software issues, but they didn't have the funding to blow up dozens of prototypes.

Carmack drove that project, he wrote code, he built rocket engines himself, he ran missions.

Elon, notably, doesn't do any of that. He just had more money, or was willing to commit more money, to seeing it through. For Carmack it was more of a fun diversion (the X Prize) than a business he wanted to build.

Post reply on HN