Live data from Hacker News

Tesla Dojo Custom AI Supercomputer at HC34

servethehome.com

11–20 of 102 posts

Re: Tesla Dojo Custom AI Supercomputer at HC34

#11
post #2

They recycled a lot of material from last year's presentation. Tesla is the only company you really have to say this about: there is a non-negligible probability that this thing doesn't exist. There are no published results and we aren't seeing the mind-blowing pace of FSD improvements that Musk promised us in 2020 when Dojo 1.0 was only a year away.

Musk promising FSD by the end of the year - year after year - is one major thing that tarnished the Tesla image in my mind. I now assume that he has to believe that in order to avoid lawsuits. Maybe the same is with the computing infrastructure. They need to build it to show that they honestly believed it could work - even though they failed to deliver on their promises many times. That said, they do make progress an…

Musk is absolutely the reason I would never consider buying another Tesla, having just sold mine.

Dude just straight up lies. He may not realize he's lying, but he lies constantly.

Edit, because I'm getting downvoted. Here's some examples: self driving, battery swaps, robotic snake chargers, cybertruck windows, Bitcoin won't be converted to fiat, starlink speeds will improve, he will sell his home, first mars mission 2024, Tesla solar roofs, power packs at every supercharger, brake pads on Tesla cars will never need to be replaced, the gigafactory will be 100% renewable powered by 2020, fixing the Flint water crisis, making bricks from boring company waste, founding a media credibility organization, and probably a TON more.

Oh man, I forgot he took money to fly tourists to the moon.

There are so many lies he's told. So many.

Re: Tesla Dojo Custom AI Supercomputer at HC34

#12
If this stuff is any good, Tesla should make it available for everyone to use colab-style.

A hosted Jupyter notebook in a sandboxed VM able to send jobs to this new silicon is something that might be possible to set up by a small team in a few months, and could turn into a billion dollar business.

As a bonus, Tesla can use revenue from that to grow this supercomputer, while using any spare/unsold capacity for themselves.

Re: Tesla Dojo Custom AI Supercomputer at HC34

#14

Earlier quoted context omitted.

Musk promising FSD by the end of the year - year after year - is one major thing that tarnished the Tesla image in my mind. I now assume that he has to believe that in order to avoid lawsuits. Maybe the same is with the computing infrastructure. They need to build it to show that they honestly believed it could work - even though they failed to deliver on their promises many times. That said, they do make progress an…

Musk is absolutely the reason I would never consider buying another Tesla, having just sold mine. Dude just straight up lies. He may not realize he's lying, but he lies constantly. Edit, because I'm getting downvoted. Here's some examples: self driving, battery swaps, robotic snake chargers, cybertruck windows, Bitcoin won't be converted to fiat, starlink speeds will improve, he will sell his home, first mars mission…

My autopilot drives me on the highway just like he said it would. That was what he predicted about 5 years ago. FSD is taking longer - so what. It's progressing for the public far more than another other attempt.

Re: Tesla Dojo Custom AI Supercomputer at HC34

#15

If this stuff is any good, Tesla should make it available for everyone to use colab-style. A hosted Jupyter notebook in a sandboxed VM able to send jobs to this new silicon is something that might be possible to set up by a small team in a few months, and could turn into a billion dollar business. As a bonus, Tesla can use revenue from that to grow this supercomputer, while using any spare/unsold capacity for themsel…

It is already Tesla's plan to build AWS-style paid access to Dojo.

I think they said that during the first AI Day. Here's a 19 minute supercut of AI Day: https://www.youtube.com/watch?v=keWEE9FwS9o

Re: Tesla Dojo Custom AI Supercomputer at HC34

#16

If this stuff is any good, Tesla should make it available for everyone to use colab-style. A hosted Jupyter notebook in a sandboxed VM able to send jobs to this new silicon is something that might be possible to set up by a small team in a few months, and could turn into a billion dollar business. As a bonus, Tesla can use revenue from that to grow this supercomputer, while using any spare/unsold capacity for themsel…

I’m almost certain that Google/Amazon/Apple et al have similar specialized computing hardware, but at smaller scales. We don’t publicly know much about their internal hardware. If they had the infra, I think it would be interesting to see some Tesla cloud service offering indeed.

[1] https://cloud.google.com/blog/products/ai-machine-learning/g...

Re: Tesla Dojo Custom AI Supercomputer at HC34

#17
post #2

They recycled a lot of material from last year's presentation. Tesla is the only company you really have to say this about: there is a non-negligible probability that this thing doesn't exist. There are no published results and we aren't seeing the mind-blowing pace of FSD improvements that Musk promised us in 2020 when Dojo 1.0 was only a year away.

Musk promising FSD by the end of the year - year after year - is one major thing that tarnished the Tesla image in my mind. I now assume that he has to believe that in order to avoid lawsuits. Maybe the same is with the computing infrastructure. They need to build it to show that they honestly believed it could work - even though they failed to deliver on their promises many times. That said, they do make progress an…

+ the robot guy dancing in spandex

Re: Tesla Dojo Custom AI Supercomputer at HC34

#18

Earlier quoted context omitted.

Musk promising FSD by the end of the year - year after year - is one major thing that tarnished the Tesla image in my mind. I now assume that he has to believe that in order to avoid lawsuits. Maybe the same is with the computing infrastructure. They need to build it to show that they honestly believed it could work - even though they failed to deliver on their promises many times. That said, they do make progress an…

Musk is absolutely the reason I would never consider buying another Tesla, having just sold mine. Dude just straight up lies. He may not realize he's lying, but he lies constantly. Edit, because I'm getting downvoted. Here's some examples: self driving, battery swaps, robotic snake chargers, cybertruck windows, Bitcoin won't be converted to fiat, starlink speeds will improve, he will sell his home, first mars mission…

No post body was provided.

Re: Tesla Dojo Custom AI Supercomputer at HC34

#19
post #9

Cool. Can anyone chime in on how this compares to other ML SoC/ASICs? I know many places are going hard on general purpose GPUs, but I’d imagine ASIC based supercomputers (like Google TPU) are the way to go forward.

They'll be at a process disadvantage. D1 is allegedly TSMC 7nm, as per last year's information (https://www.tomshardware.com/news/tesla-d1-ai-chip).

An ASIC can strip out features they don't need and save some space. But a good chunk of modern GPUs are memory-controllers, registers, and SIMD-cores. And modern GPUs (both AMD's MI250x and NVidia's A100) have 16-bit matrix multiplication units (aka: Tensor cores). Once we factor in the process disadvantage, I'm not sure if the D1 will be as competitive as they hope.

Tesla's hope is that their D1 chip has more 16-bit matrix multiplication cores than the NVidia / AMD designs. But A100 is quite solid, and NVidia Hopper has been announced at HC34 (aka: NVidia's next generation).

https://www.nvidia.com/en-us/technologies/hopper-architectur...

-------

Most of this presentation on the Tesla Dojo is about the interconnect system. Alas, NVidia's on like the 4th (or was it 5th?) generation of NVlink, available from their DGX servers (and I'm sure a Hopper version will come out soon).

AMD's not far behind, also with a lot of good presentations this year from HC34 that points out how AMD's "Frontier" Supercomputer has huge bandwidth. In particular, each MI250X GPU is a twin-chiplet design (two GPUs per... GPU), with 5 high-speed links to connect to other GPUs in a high-speed fashion. There's a reason why Frontier is the #1 supercomputer of the world right now, in both absolute Double-precision FLOPs, and in Green Double-precision FLOPS-per-watt.

NVidia's Hopper will be hitting 4nm. AMD's MI250x is 5nm. That means the D1 chip has less than 1/2 the transistors at the same area compared to NVidia.

> but I’d imagine ASIC based supercomputers (like Google TPU) are the way to go forward.

Only if you keep up with the process shrinks. 7nm is getting long in the tooth now. All eyes are looking forward to 5nm, 4nm, and even 3nm designs (now that Apple is the customer of TSMC's 3nm node).

-----------

That being said, if the 7nm node is cheaper, maybe this exercise was still cost-effective for Tesla. As the newer nodes obsolete the older nodes, the older nodes become more cost-effective.

Cost-efficiency is less popular / less cool, but still an effective business plan.

> but I’d imagine ASIC based supercomputers (like Google TPU) are the way to go forward.

The issue is that it probably costs hundreds-of-millions of dollars to design something like the D1. Sure, the mass production of the chip afterwards will be incredible, but chips have stupidly high startup costs (masking, engineering, research, etc. etc.)

GPUs on the other hand, are more general purpose and are applicable to more situations. So you can sell the GPU to more customers and spread out the R&D costs. In particular, GPUs capture the attention of the video game crowd, who will fund high-end GPU research just to play video games.

Much like how Intel's laptops allow servers to share the R&D effort, so too does NVidia's consumer GPUs share research/development costs with their high end A100 cards.

Re: Tesla Dojo Custom AI Supercomputer at HC34

#20
post #16

If this stuff is any good, Tesla should make it available for everyone to use colab-style. A hosted Jupyter notebook in a sandboxed VM able to send jobs to this new silicon is something that might be possible to set up by a small team in a few months, and could turn into a billion dollar business. As a bonus, Tesla can use revenue from that to grow this supercomputer, while using any spare/unsold capacity for themsel…

I’m almost certain that Google/Amazon/Apple et al have similar specialized computing hardware, but at smaller scales. We don’t publicly know much about their internal hardware. If they had the infra, I think it would be interesting to see some Tesla cloud service offering indeed. [1] https://cloud.google.com/blog/products/ai-machine-learning/g...

Google has had TPUs for a while. Most of the incredibly powerful transformer architectures like Parti and PaLM are trained on it. They even merge TPU pods with Pathways so they can go up to 540B parameters (gpt-3 is only 175B)

https://blog.google/technology/ai/introducing-pathways-next-...

Post reply on HN