Live data from Hacker News

Tesla turns on 10k-node Nvidia H100 Cluster

techradar.com

91–100 of 135 posts

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#91
post #83
post #41

Earlier quoted context omitted.

> What makes you think that Telsa, a company with far less AI workers and knowledge, an far less money than the above companies can out design and out build them? Presumably because Elon himself will be involved in the design, and Elon, as we all know, is one of the world's great thinkers. ;)

Elon is one of the worlds greatest talent poachers, and that is much better than being a great thinker.

Do you mind explaining what makes him good at it? Pay? Atmosphere? Management style?

From what I heard about SpaceX it seems to be a place grads go to burn out while being paid below market rate simply because they're excited about the idea. Maybe that impression is wrong, so I'd like to hear other perspectives.

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#92

Earlier quoted context omitted.

Linux was a lot harder back then.

Harder for who? Elon certainly didn't have the technical chops to work with it.

Harder for everyone, including his staff, who were asking him to move to Windows…

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#93

I understand that the H100 is NVidia's leading edge chip, but can someone let me know if 10K is considered to be a big cluster? I've never worked inside one of the leading edge AI companies like OpenAI, Google, Microsoft or Meta. Is this comparable to what they would work with? My first guess is that it seems much smaller. And if you are running many parallel training jobs then you are getting about 1,000 chips at mo…

Last I heard, the estimate was that NVIDIA would build 550k units in 2023, so 2% of all production — especially as at least six others (your four plus Apple and at least one intelligence agency) will be of similar size by themselves — is certainly non-negligible.

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#95
post #83
post #41

Earlier quoted context omitted.

> What makes you think that Telsa, a company with far less AI workers and knowledge, an far less money than the above companies can out design and out build them? Presumably because Elon himself will be involved in the design, and Elon, as we all know, is one of the world's great thinkers. ;)

Elon is one of the worlds greatest talent poachers, and that is much better than being a great thinker.

Is he, though?

I recently spoke with someone who quit SpaceX because (among other reasons) they felt Elon was a meddling micro-manager. That's just one anecdote, of course, but the Internet is full of them, replete with summary firings (https://www.businessinsider.com/tesla-elon-musk-ruthlessly-f...), worker safety issues (https://www.washingtonpost.com/technology/2021/03/12/hundred...), and just general bullshit (https://www.reddit.com/r/EnoughMuskSpam/comments/9e360m/elon...).

I don't deny that his public image, for years, was an overall positive one. I really enjoyed Jill Lepore's digging into it here: https://www.pushkin.fm/podcasts/elon-musk-the-evening-rocket.

But it seems like people who worked with him knew, for a long time, that he was full of shit. And increasingly, the public seems to as well.

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#96
post #60

Earlier quoted context omitted.

In the old days, depending on architecture, fp64 performance could be atrocious even when fp32 was decent, so bragging about fp64 performance has an authenticity to it. Not all scientific computing requires 64 bits, but knowing that you can drop to high precision when necessary without penalty is nice. Also, back in the day, integer ops were just called 'ops', grumble grumble. But yeah FLOPS specifically refers to fl…

Still true that fp64 throughput is lower for consumer GPUs - both NV and AMD. That’s kinda why I was curious about leading with that metric - outside of HPC and scientific applications, a lot of people don’t really need fp64, and the machine might normally have a much higher fp32 throughput. > knowing you can drop to high precision when necessary without penalty is nice. I guess I maybe don’t know why you’d ever have…

This is also related to "fast" versions of all some operations. You might want the full 32 bit float but you dont want or need to do full precision division or sqrt operations. This is common in games/graphics and probably machine learning.

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#98
post #49

Earlier quoted context omitted.

He delivered on them though right? Also reusable rockets didn’t exist?

John Carmack on a shoestring budget nearly got this working at Armadillo. If he had more money and time he would've had it working a half decade before SpaceX.

"Nearly" and "if he had more money and time" is "no".

And given how much faster SpaceX has been than anyone else, I can only believe this "could've, would've" in form of the slightly longer hypothetical "if only Carmack hired all the rocket scientists (and raised all the money to give them freedom to go fast) before Musk got there".

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#99
post #75
post #71

Earlier quoted context omitted.

> high editorial standards of such tech-press luminaries as "TechRadar" and "Hacker News". If you would’ve just scrolled just a little bit on that Twitter post that you linked. You would’ve seen these: https://x.com/sawyermerritt/status/1696012091964915744 https://x.com/tim_zaman/status/1695488119729238147 Also, just FYI. Sawyer posts most of the Tesla and SpaceX breaking news on Twitter before major outlets even wri…

> If you would’ve just scrolled just a little bit on that Twitter post that you linked. You would’ve seen these: I don't see those when I scroll. I see "Buckle up everyone, the acceleration of progress is about to get nutty!" and this is the end of the post? Maybe I'm misusing this thing? > https://x.com/tim_zaman/status/1695488119729238147 So another guy who claims to be a Tesla employee says (again, strangely futur…

> Maybe I'm misusing this thing?

That seems to be the case here. ;)

> So another guy who claims to be a Tesla employee says (again, strangely future tense) that this is true? I mean, I am willing to believe--'cause he paid $20 for a blue check--that he probably is a Tesla employee.

Another case of misuse? Here’s a tip for you. When you see a company logo/icon on someone's Twitter/X profile. That means they are verified to be affiliated with that org.

“Accounts affiliated with the organization will receive an affiliate badge on their profile with the organization’s logo, and will be featured on the organization’s Twitter profile, indicating their affiliation. “

https://twitter.com/verified/status/1641596848921276417

Instead of inferring that Tim Zaman is a random Twitter user who paid $20 for a blue check. Why not just Google his name? ;)

https://letmegooglethat.com/?q=Tim+Zaman

> I guess I'm old. Back in my day, "evidence" wasn't some random dude's online posts. But I know things have changed. ;)

I linked a video where CNBC was interviewing Sawyer but it seems that you didn’t even bother to check it.

This seems to be the problem today. People refuse to do the bare minimum (which is not even much) required for critical thinking. Instead of verifying information, people tend to uncritically repeat inaccurate assumptions, even when provided with additional information in good faith.

Re: Tesla turns on 10k-node Nvidia H100 Cluster

#100

Earlier quoted context omitted.

Harder for who? Elon certainly didn't have the technical chops to work with it.

Harder for everyone, including his staff, who were asking him to move to Windows…

He should have hired staff that is competent with the tech stack used at his company.

Unforced rewrites are usually always a bad idea.

Post reply on HN