Live data from Hacker News

Tesla turns on 10k-node Nvidia H100 Cluster

techradar.com

41–50 of 135 posts

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#41

Earlier quoted context omitted.

It’s bottleneck on Nvidia side. They are producing less than Tesla consume. Tesla compute power will outclass many cloud provider combined in just three or four years with their own custom chip.

> It’s bottleneck on Nvidia side. They are producing less than Tesla consume. Tesla compute power will outclass many cloud provider combined in just three or four years with their own custom chip. That seems like a bold claim. Google, Microsoft and Meta make so much more money than Telsa that if making AI chips was so easy, then they could clearly out design and build Tesla without thinking too hard about it. What ma…

> What makes you think that Telsa, a company with far less AI workers and knowledge, an far less money than the above companies can out design and out build them?

Presumably because Elon himself will be involved in the design, and Elon, as we all know, is one of the world's great thinkers. ;)

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#42
post #7

> The firm also built a compute cluster fitted with 5,760 Nvidia A100 GPUs in June 2012 Wow, that's some really early hardware access. /s

Maybe, they picked up date when Elon first communicated that they are "ready" to go live. Like everything else it took a decade to materialize.

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#43
post #12

What happened to their custom hardware training stack Dojo? They had some interesting ideas there. The last I heard, they had one of those tiles "working" in the lab. Pretty far from a production setup. I can imagine they either underestimated the software effort needed to squeeze as much performance as possible out of those things, or they underestimated the pace at which Nvidia scales FLOPS/$, or both.

[deleted]

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#44
post #35
post #7

> The firm also built a compute cluster fitted with 5,760 Nvidia A100 GPUs in June 2012 Wow, that's some really early hardware access. /s

Lol, I was wondering if A100 is really that old. Turns out A100 was released in 2020.

Yea I assume they meant 2021. 2012 was still the early days of GPU compute. Best we had were M2090s.

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#45
post #29

I understand that the H100 is NVidia's leading edge chip, but can someone let me know if 10K is considered to be a big cluster? I've never worked inside one of the leading edge AI companies like OpenAI, Google, Microsoft or Meta. Is this comparable to what they would work with? My first guess is that it seems much smaller. And if you are running many parallel training jobs then you are getting about 1,000 chips at mo…

The most powerful listed supercomputer has 37,888 Radeon GPUs, so in the same order of magnitude.

Interesting choice of words... I take you work for OpenAI? :) How large is their/'your' cluster? Probably the biggest in the world by now..

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#46
post #12

What happened to their custom hardware training stack Dojo? They had some interesting ideas there. The last I heard, they had one of those tiles "working" in the lab. Pretty far from a production setup. I can imagine they either underestimated the software effort needed to squeeze as much performance as possible out of those things, or they underestimated the pace at which Nvidia scales FLOPS/$, or both.

Vaporware, just like much of what Musk talks about.

What that he has talked about been vaporware?

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#47
post #22

> This AI cluster, worth more than $300 million, will offer a peak performance of 340 FP64 PFLOPS for technical computing and 39.58 INT8 ExaFLOPS for AI applications, according to Tom’s Hardware. I was curious why this statement lead with fp64 flops (instead of fp32, perhaps), but I looked up the H100 specs, and NV’s marketing page does the same thing. They’re obviously talking about the H100 SXM here, which has the…

In the old days, depending on architecture, fp64 performance could be atrocious even when fp32 was decent, so bragging about fp64 performance has an authenticity to it. Not all scientific computing requires 64 bits, but knowing that you can drop to high precision when necessary without penalty is nice.

Also, back in the day, integer ops were just called 'ops', grumble grumble. But yeah FLOPS specifically refers to floating point. Calling them TOPS doesn't make sense to me, since tensor cores were meant for matrix operation speedup, and these matrices are rarely integer.

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#48
post #5

Original tweet: https://twitter.com/SawyerMerritt/status/1696011140508045660 Previus article: https://www.tomshardware.com/news/teslas-dollar300-million-a... This is second-hand blogspam.

Tom's Hardware and Tech Radar belong to the same company. If you consider this to be blog spam, almost any news website these days would be blog spam.

> almost any news website these days would be blog spam

Yes.

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#49

Earlier quoted context omitted.

Reusable rockets, electric cars, solar panels... What would you say grants you the standing to opine here?

I'm fairly certain all of those existed prior to Musk's suggestion of them.

He delivered on them though right? Also reusable rockets didn’t exist?

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#50
I'm confused. The article from September 1 linked to here is strangely future-tense ("But the firm’s latest investment in 10,000 of the company’s H100 GPUs dwarfs the power of this supercomputer....This AI cluster, worth more than $300 million, will offer a peak performance...").

It links to a Tom's Hardware article (https://www.tomshardware.com/news/teslas-dollar300-million-a...) from August 28 that says "Tesla is about to flip the switch on its new AI cluster, featuring 10,000 Nvidia H100 compute GPUs") and says "Tesla is set to launch its highly-anticipated supercomputer on Monday..." (presumably the September 1 event).

So, like, does Tesla actually have 10k H100s? Or do they have an order for 10k H100s? Or an intention to buy 10k H100s?

Is the sole source for these articles this (https://twitter.com/SawyerMerritt/status/1696011140508045660) random Twitter post by some guy who runs an online clothing company?

I don't mean to snipe, but this article doesn't seem to rise to the extremely high editorial standards of such tech-press luminaries as "TechRadar" and "Hacker News".

Post reply on HN