Live data from Hacker News

An inside look at the custom CPUs in Tesla's Dojo Supercomputer

semianalysis.com

121–130 of 132 posts

Re: An inside look at the custom CPUs in Tesla's Dojo Supercomputer

#121

Earlier quoted context omitted.

There’s no third party here. As far as I can tell Tesla designed the chips and wrote the software stack to work together. They don’t want to rely on third parties for their critical infrastructure and AI chip design is vital to their success. At least this is what I can grok as an outsider.

Right, I understand the claims, but as you can see with AMD there's a ton of work to actually be competitive, and AMD has a large team working on it. I'm just skeptical at the claim the the software is of any use to anyone but a small group.

I think the difference is that AMD needs to cater to a wide variety of customers, software stacks, and hardware platforms. Tesla has basically one platform and one customer (themselves).

Their software only needs to be useful to a small group - the autonomy team. I did not get the impression that Tesla plans to sell supercomputers, but that they are building these supercomputers for themselves to train and deploy AI networks.

You said "Even if the hardware existed, if you can't program it, it won't succeed." But it seems like the key customer - Tesla's internal Autopilot team - can already program it. So I just don't see any problem. They may choose to sell these systems (note that they have been shipping their Gen1 custom chip for years now and it is not for sale to the public), but the real plan for revenue is to succeed at AI and profit there. For that they need incredible computing power internally and in a portable platform for edge deployment, but they do not need to sell their chips as a general purpose compute platform.

Re: An inside look at the custom CPUs in Tesla's Dojo Supercomputer

#122
post #69

So what's the programming model, ISA and compiler technology situation (reliance on / maturity of) like?

replying to myself: according to the transcript[1] it's a custom ISA and they mentioned some things about their custom compiler, including a diagram[2].

[1]https://www.tweaktown.com/news/81229/teslas-insane-new-dojo-...

[2] https://www.tweaktown.com/image.php?image=https://static.twe...

Re: An inside look at the custom CPUs in Tesla's Dojo Supercomputer

#123

Earlier quoted context omitted.

Right, I understand the claims, but as you can see with AMD there's a ton of work to actually be competitive, and AMD has a large team working on it. I'm just skeptical at the claim the the software is of any use to anyone but a small group.

I think the difference is that AMD needs to cater to a wide variety of customers, software stacks, and hardware platforms. Tesla has basically one platform and one customer (themselves). Their software only needs to be useful to a small group - the autonomy team. I did not get the impression that Tesla plans to sell supercomputers, but that they are building these supercomputers for themselves to train and deploy AI…

My problem with articles like these is the sensational headlines that Tesla made an AMD/Nvidia killer. AMD and Nvidia could both make this chip if they wanted to, but dedicating all the die to matrix multiply cores is a waste for most end users. The media makes it seem like Tesla did something revolutionary here, but all they did was make a very, very targeted asic.

Re: An inside look at the custom CPUs in Tesla's Dojo Supercomputer

#124

Earlier quoted context omitted.

I think the difference is that AMD needs to cater to a wide variety of customers, software stacks, and hardware platforms. Tesla has basically one platform and one customer (themselves). Their software only needs to be useful to a small group - the autonomy team. I did not get the impression that Tesla plans to sell supercomputers, but that they are building these supercomputers for themselves to train and deploy AI…

My problem with articles like these is the sensational headlines that Tesla made an AMD/Nvidia killer. AMD and Nvidia could both make this chip if they wanted to, but dedicating all the die to matrix multiply cores is a waste for most end users. The media makes it seem like Tesla did something revolutionary here, but all they did was make a very, very targeted asic.

That makes sense. Certainly lots of journalism is bad. I haven’t read the article actually. I’ve just watched the two hour presentation by Tesla as well as their past presentations. I am a robotics engineer and I’ve been trying to understand how best to make an “animal like” brain system for an autonomous robot in the real world. I have been pleased with how much Tesla shared about their system and I think their extremely powerful hardware and their neural network approach is ideal for solving this problem. So I’m very happy with what they’ve come up with and I’m happy that it will push competitors to do the same as I think very large neural networks might be needed to solve general purpose robotics.

So all in all I think the chip is very good and I think they are on a path to success. Whatever the article says to hype it up doesn’t change the value of what Tesla has done in my eyes.

Re: An inside look at the custom CPUs in Tesla's Dojo Supercomputer

#125
post #75

The fact that they didn't do this: > Their current training cluster would be the 5th largest supercomputer if Tesla stopped all real workloads, ran Linpack, and submitted it to the Top500 list. which is trivial to do, and pretty much a must when bringing up the cluster to make sure its working properly, so much that most clusters do this on every maintainance, along with another bunch of benchmarks; and that they say…

> Their current training cluster would be the 5th largest supercomputer if Tesla stopped all real workloads, ran Linpack, and submitted it to the Top500 list.

That line is referring to their current Nvidia A100 powered supercomputer which they set up with 5,760 A100 GPUs recently this year [0].

Read the previous line of the line which you’ve shared from the post:

> Tesla has been expanding the size of their GPU clusters for years. Their current training cluster would be the 5th largest supercomputer if Tesla stopped all real workloads, ran Linpack, and submitted it to the Top500 list.

[0] https://blogs.nvidia.com/blog/2021/06/22/tesla-av-training-s...

Re: An inside look at the custom CPUs in Tesla's Dojo Supercomputer

#126
post #61

Earlier quoted context omitted.

If you look at CPU performance relative to density as it has progressed over the last few decades, there's a clear decline in speed improvement. Denser only means faster to a point, regardless of Intel's process update failures.

Not sure what you’re replying to. The statement I’m replying to didn’t say anything about performance. It said “ Transistor density isn’t doubling.”

I have always looked at Moore's "law" as a proxy for performance, so I felt it was relevant, if tangential.

Re: An inside look at the custom CPUs in Tesla's Dojo Supercomputer

#127

i cant stop thinking about the tesla bot. was he he outright lying? is it just a ploy to recruit robotics people? is it even plausible? i think the most challenging aspect of the idea is interacting with the world, picking up and handling various objects. its obvious from the presentation that fsd is very good at placing itself in space and mapping out its environment as well as devising routes even when accounting f…

>is it even plausible? The last commercial anthropomorphic was the Willow Garage PR2 back in 2010. It weighed 600 pounds, and had a wheeled base. Each arm had a max payload of 4 pounds. It cost $250,000. The company went bankrupt because there wasn't anything you could do with it. The tesla bot is supposed to be bipedal, only weigh 125 pounds, and have a "arm extend lift" of 10 lbs. Is that per arm, or both together?…

> Either onboard compute is minimal, or it has a battery life measured in dozens of minutes.

For industrial applications you could probably have some sort of novel power system, like a tether from the ceiling or special floor that delivers power through the feet.

Re: An inside look at the custom CPUs in Tesla's Dojo Supercomputer

#128

Earlier quoted context omitted.

I think the difference is that AMD needs to cater to a wide variety of customers, software stacks, and hardware platforms. Tesla has basically one platform and one customer (themselves). Their software only needs to be useful to a small group - the autonomy team. I did not get the impression that Tesla plans to sell supercomputers, but that they are building these supercomputers for themselves to train and deploy AI…

My problem with articles like these is the sensational headlines that Tesla made an AMD/Nvidia killer. AMD and Nvidia could both make this chip if they wanted to, but dedicating all the die to matrix multiply cores is a waste for most end users. The media makes it seem like Tesla did something revolutionary here, but all they did was make a very, very targeted asic.

If anything I think they would 'sell' this as a cloud service I would think. But it would be a while before that happens.

Re: An inside look at the custom CPUs in Tesla's Dojo Supercomputer

#129
post #94

Earlier quoted context omitted.

In the presentation they were quite open that they got to the point of running real loads but only on a single tile on a bench. Not clear what you're disputing here?

I only work on a lowly private cluster but running standard benchmarks is utterly routine here (it is in fact automated). As others with HPC experience pointed out, running benchmarks is pretty much mandatory when bringing a new system up, not just to ensure it actually performs as promised, but also to weed out bad and marginal components. You do one or two weeks of intense benchmarking and testing and you're assure…

> So it is literally unbelievable that Tesla not just stands up a cluster, but created their own hardware to do so, and didn't run any quotable benchmark and only has the theoretic FLOPS numbers for marketing.

They haven't gotten to this point yet.

They have a single tile (A node). It wouldn't surprise me if the tile they have is just a prototype as well.

You have to have a cluster before you can start running cluster benchmarks.

Re: An inside look at the custom CPUs in Tesla's Dojo Supercomputer

#130
post #75

The fact that they didn't do this: > Their current training cluster would be the 5th largest supercomputer if Tesla stopped all real workloads, ran Linpack, and submitted it to the Top500 list. which is trivial to do, and pretty much a must when bringing up the cluster to make sure its working properly, so much that most clusters do this on every maintainance, along with another bunch of benchmarks; and that they say…

Had to look into MLPerf as I don’t follow super compute.

But I don’t see them not doing that test as an issue as they made it quite clear that their entire system is tailor made, like an ASIC, to focus on neural nets and specific, relevant compute pipelines to what they care about. I would imagine dropping a generic, broad based ML benchmarking tool will not only perform suboptimally but also not be representative of what they’re trying to do. It’s not meant to be a general purpose ML super computer, it’s supposed to be a super computer to solve a narrowish niche of problems.

Post reply on HN