Live data from Hacker News

Tesla turns on 10k-node Nvidia H100 Cluster

techradar.com

71–80 of 135 posts

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#71
post #50

I'm confused. The article from September 1 linked to here is strangely future-tense ("But the firm’s latest investment in 10,000 of the company’s H100 GPUs dwarfs the power of this supercomputer....This AI cluster, worth more than $300 million, will offer a peak performance..."). It links to a Tom's Hardware article ( https://www.tomshardware.com/news/teslas-dollar300-million-a... ) from August 28 that says "Tesla is…

> high editorial standards of such tech-press luminaries as "TechRadar" and "Hacker News".

If you would’ve just scrolled just a little bit on that Twitter post that you linked. You would’ve seen these:

https://x.com/sawyermerritt/status/1696012091964915744

https://x.com/tim_zaman/status/1695488119729238147

Also, just FYI. Sawyer posts most of the Tesla and SpaceX breaking news on Twitter before major outlets even write their articles.

For example, here’s one just 12mins ago as confirmed by Elon: https://x.com/sawyermerritt/status/1728092021628313777

A “random Twitter post by some guy who runs an online clothing company” is definitely a wrong assumption.

https://x.com/sawyermerritt/status/1709019899442479162

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#72

Earlier quoted context omitted.

Reusable rockets, electric cars, solar panels... What would you say grants you the standing to opine here?

I'm fairly certain all of those existed prior to Musk's suggestion of them.

[dead]

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#74

Earlier quoted context omitted.

It’s bottleneck on Nvidia side. They are producing less than Tesla consume. Tesla compute power will outclass many cloud provider combined in just three or four years with their own custom chip.

> It’s bottleneck on Nvidia side. They are producing less than Tesla consume. Tesla compute power will outclass many cloud provider combined in just three or four years with their own custom chip. That seems like a bold claim. Google, Microsoft and Meta make so much more money than Telsa that if making AI chips was so easy, then they could clearly out design and build Tesla without thinking too hard about it. What ma…

The dirty little open secret with a lot of these platforms is the contract sizes, hardware costs, etc are so massive they come with multiple teams of dedicated engineers and internal expertise to get your application(s) up and running on them. Obviously these things are never quite "pull a docker container and run" and no one dropping eight-nine figures on these installs is going to do it without serious vendor backing and support.

It's part of the reason why AMD has had quite a bit of success here but is in single digit market share for "AI" otherwise.

Most people - even large orgs with thousands of GPUs - are so trapped in CUDA the theoretical on paper performance and cost benefits evaporate immediately when you spend all of your time trying to port everything over to the point you get equivalent performance and functionality.

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#75
post #71
post #50

I'm confused. The article from September 1 linked to here is strangely future-tense ("But the firm’s latest investment in 10,000 of the company’s H100 GPUs dwarfs the power of this supercomputer....This AI cluster, worth more than $300 million, will offer a peak performance..."). It links to a Tom's Hardware article ( https://www.tomshardware.com/news/teslas-dollar300-million-a... ) from August 28 that says "Tesla is…

> high editorial standards of such tech-press luminaries as "TechRadar" and "Hacker News". If you would’ve just scrolled just a little bit on that Twitter post that you linked. You would’ve seen these: https://x.com/sawyermerritt/status/1696012091964915744 https://x.com/tim_zaman/status/1695488119729238147 Also, just FYI. Sawyer posts most of the Tesla and SpaceX breaking news on Twitter before major outlets even wri…

> If you would’ve just scrolled just a little bit on that Twitter post that you linked. You would’ve seen these:

I don't see those when I scroll. I see

"Buckle up everyone, the acceleration of progress is about to get nutty!"

and this is the end of the post?

Maybe I'm misusing this thing?

> https://x.com/tim_zaman/status/1695488119729238147

So another guy who claims to be a Tesla employee says (again, strangely future tense) that this is true? I mean, I am willing to believe--'cause he paid $20 for a blue check--that he probably is a Tesla employee.

But the use of future tense is a bit weird, right? And the lack of any followup?

> A “random Twitter post by some guy who runs an online clothing company” is definitely a wrong assumption.

I guess I'm old. Back in my day, "evidence" wasn't some random dude's online posts. But I know things have changed. ;)

==

More seriously:

https://www.hpcwire.com/2023/08/17/nvidia-h100-are-550000-gp... says Nvidia is producing 550k H100s in 2023. And there's obviously a significant lead-time requirement.

So, yes, I can sorta imagine Tesla pre-ordered 2% of global supply of H100s early in 2023 and was bragging about it at the end of August just 'cause.

But I can also imagine this is smoke and mirrors, and they have, like, a handful with the rest on backorder, and we haven't heard more about it 'cause Tesla doesn't have marketing people, it just has wahoos who post things on Twitter.

Either way, I guess?

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#76
post #46

Earlier quoted context omitted.

What that he has talked about been vaporware?

https://elonmusk.today/

Hmm that website could be really interesting if it clearly didn’t try being misleading, e.g. a statement about something being in development but not yet out isn’t a promise that it will be out now. Really silly.

Also if it didn’t exclude everything that has been delivered. That long list would be interesting to see as well.

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#77

I understand that the H100 is NVidia's leading edge chip, but can someone let me know if 10K is considered to be a big cluster? I've never worked inside one of the leading edge AI companies like OpenAI, Google, Microsoft or Meta. Is this comparable to what they would work with? My first guess is that it seems much smaller. And if you are running many parallel training jobs then you are getting about 1,000 chips at mo…

I previously ran 150,000 AMD GPUs. 10k doesn't seem that large. =)

That said, these GPUs aren't just the GPUs. They are whole chassis. They are huge onboard storage arrays, TB's of RAM, 800G networking (and associated cables), racks, cooling, power distribution, backup power, etc...

None of it is easy.

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#78

I understand that the H100 is NVidia's leading edge chip, but can someone let me know if 10K is considered to be a big cluster? I've never worked inside one of the leading edge AI companies like OpenAI, Google, Microsoft or Meta. Is this comparable to what they would work with? My first guess is that it seems much smaller. And if you are running many parallel training jobs then you are getting about 1,000 chips at mo…

This is a big cluster, definitely large enough to pretrain 100B+ parameter LLMs in months. Source - I work at Databricks in the ML platform.

I don’t know much about AV processing, that’s highly customized to only a few customers but I’d expect it to also have very large computational requirements to do video processing and reinforcement learning.

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#80

I understand that the H100 is NVidia's leading edge chip, but can someone let me know if 10K is considered to be a big cluster? I've never worked inside one of the leading edge AI companies like OpenAI, Google, Microsoft or Meta. Is this comparable to what they would work with? My first guess is that it seems much smaller. And if you are running many parallel training jobs then you are getting about 1,000 chips at mo…

I previously ran 150,000 AMD GPUs. 10k doesn't seem that large. =) That said, these GPUs aren't just the GPUs. They are whole chassis. They are huge onboard storage arrays, TB's of RAM, 800G networking (and associated cables), racks, cooling, power distribution, backup power, etc... None of it is easy.

Out of interest, what did you use all that compute for?
Post reply on HN